A workspace for better LLM decisions

Know which LLM setup works best.

Compare models, develop prompts, and evaluate outputs in one workspace—so you can make decisions with evidence.

Model benchmarksVersioned promptsReproducible evaluations

One workspace. Three ways to improve.

Explore real screens. No sign-in needed for this walkthrough.

Read-only examples
LLMForge benchmark comparison showing quality and latency for two seeded example configurationsLLMForge benchmark comparison showing quality and latency for two seeded example configurations

Compare quality, latency and usage together.

Inspect quality, latency and token usage side by side on the same test set. Keep the configuration and per-case outputs behind every result.

Seeded examples: A is Compact setup; B is Deliberate setup. These illustrative measurements in the real workspace are not claims about model performance.

A clear path to the next improvement.

Benchmarks compare model setups. Evaluations test saved prompt versions. Both help you decide what to change next.

  1. 01

    Configure

    Choose models for a benchmark, or save a prompt version and a fixed set of test cases.

  2. 02

    Compare

    Inspect benchmark trade-offs or compare compatible evaluation runs. Read the outputs behind the numbers.

  3. 03

    Improve

    Use disagreements and failed checks to guide your next change. Test it again before moving a prompt release label.

The tools behind a better setup.

Keep configurations, outputs and checks close to the decisions they inform.

Find the model trade-off that fits.

Compare benchmark runs on fixed examples. Inspect accuracy, latency, token usage and disagreements before choosing a setup.

Explore the walkthrough
LLMForge benchmark comparison showing quality and latency for two seeded example configurationsLLMForge benchmark comparison showing quality and latency for two seeded example configurations
Seeded examples: A is Compact setup; B is Deliberate setup. These illustrative measurements in the real workspace are not claims about model performance.

Turn a prompt draft into a tested release.

Develop text or chat prompts, test candidates in the playground, and retain the version history behind each release.

Explore the walkthrough
LLMForge Echo demo prompt editor with a saved template and format controlsLLMForge Echo demo prompt editor with a saved template and format controls
The Echo demo prompt in the actual prompt editor. Demo text is deliberately simple; live playground calls require your provider key.

Catch failures before the next release.

Use assertions or versioned evaluators, inspect score coverage and reasons, and apply quality thresholds through SDK and CI gates.

Explore the walkthrough
LLMForge provider-free evaluation example with failed checks and candidate outputLLMForge provider-free evaluation example with failed checks and candidate output
A provider-free Echo demo with saved checks. It demonstrates the evaluation workflow, not real model quality.

Try the workflow. See how it works.

LLMForge is a personal engineering project, with a working demo and public implementation.

Built in the open

Inspect the code, setup instructions and benchmark methodology in the repository.

View repository

Connect your application

Fetch released prompts with the Python SDK, submit evaluation results and apply quality thresholds in CI.

Read the SDK guide

Start without model credits

The provider-free Echo demo makes prompt versions, dataset checks and comparison easy to explore.

Follow the demo

A few useful answers.

Do I need to sign in?

The walkthrough and documentation are public. Sign in to open the workspace, select a project and save prompts, datasets or runs.

Do I need a provider API key?

The Echo demo needs no model credits. Live prompt generation, judging and provider-backed model runs need the relevant provider credentials. Evaluation keys stay in memory and must be supplied again after an interruption.

Are these real model results?

These are screenshots of the actual application with deliberately prepared examples. Seeded benchmark measurements and Echo demo checks illustrate the workflow; they do not measure real model quality or establish a winning model.

How do I try the workspace demo?

Sign in, select a project and choose Create demo examples on Overview. This prepares an Echo demo prompt and a Greetings dataset. Follow the demo guide to run checks and compare results.

Can I run it locally and use it in my application?

Yes. The repository documents local setup with Next.js, FastAPI, PostgreSQL and Clerk. The Python SDK fetches released prompts and submits evaluation results; CLI quality gates can check them in CI.

Make the next change with confidence.

Sign in to your workspace, then choose Create demo examples on Overview.