01 / Benchmarks
Find the model trade-off that fits.
Compare benchmark runs on fixed examples. Inspect accuracy, latency, token usage and disagreements before choosing a setup.
Explore the walkthrough

A workspace for better LLM decisions
Compare models, develop prompts, and evaluate outputs in one workspace—so you can make decisions with evidence.
A look inside
Explore real screens. No sign-in needed for this walkthrough.


Inspect quality, latency and token usage side by side on the same test set. Keep the configuration and per-case outputs behind every result.
Seeded examples: A is Compact setup; B is Deliberate setup. These illustrative measurements in the real workspace are not claims about model performance.


Edit a candidate, keep immutable versions, and move release labels when it is ready. Your application can fetch a released prompt through the Python SDK.
The Echo demo prompt in the actual prompt editor. Demo text is deliberately simple; live playground calls require your provider key.


Evaluate a saved prompt against a fixed dataset revision. Inspect outputs and checks, compare compatible runs, and improve the cases that fail.
A provider-free Echo demo with saved checks. It demonstrates the evaluation workflow, not real model quality.
From a question to a decision
Benchmarks compare model setups. Evaluations test saved prompt versions. Both help you decide what to change next.
Choose models for a benchmark, or save a prompt version and a fixed set of test cases.
Inspect benchmark trade-offs or compare compatible evaluation runs. Read the outputs behind the numbers.
Use disagreements and failed checks to guide your next change. Test it again before moving a prompt release label.
Build with evidence
Keep configurations, outputs and checks close to the decisions they inform.
01 / Benchmarks
Compare benchmark runs on fixed examples. Inspect accuracy, latency, token usage and disagreements before choosing a setup.
Explore the walkthrough

02 / Prompts
Develop text or chat prompts, test candidates in the playground, and retain the version history behind each release.
Explore the walkthrough

03 / Evaluations
Use assertions or versioned evaluators, inspect score coverage and reasons, and apply quality thresholds through SDK and CI gates.
Explore the walkthrough

Open code. Inspectable results.
LLMForge is a personal engineering project, with a working demo and public implementation.
Inspect the code, setup instructions and benchmark methodology in the repository.
View repositoryFetch released prompts with the Python SDK, submit evaluation results and apply quality thresholds in CI.
Read the SDK guideThe provider-free Echo demo makes prompt versions, dataset checks and comparison easy to explore.
Follow the demoBefore you start
The walkthrough and documentation are public. Sign in to open the workspace, select a project and save prompts, datasets or runs.
The Echo demo needs no model credits. Live prompt generation, judging and provider-backed model runs need the relevant provider credentials. Evaluation keys stay in memory and must be supplied again after an interruption.
These are screenshots of the actual application with deliberately prepared examples. Seeded benchmark measurements and Echo demo checks illustrate the workflow; they do not measure real model quality or establish a winning model.
Sign in, select a project and choose Create demo examples on Overview. This prepares an Echo demo prompt and a Greetings dataset. Follow the demo guide to run checks and compare results.
Yes. The repository documents local setup with Next.js, FastAPI, PostgreSQL and Clerk. The Python SDK fetches released prompts and submits evaluation results; CLI quality gates can check them in CI.
Your next experiment starts here
Sign in to your workspace, then choose Create demo examples on Overview.