Developer guide

A tested prompt, from draft to application.

Keep prompt releases and test cases in one project. Inspect regressions in the workspace, fetch a released prompt in Python, and use evaluation results to gate CI.

Five-minute demo without model credits

  1. Sign in, select a project, and choose Create demo examples on Overview. This safely creates or reuses Echo demo v1 with {{query}} and a two-case Greetings dataset in this project.
  2. Open the preselected evaluation setup. The prompt version and dataset revision are fixed saved snapshots.
  3. Evaluate prompt v1 on dataset revision 1 using Demo and Exact match. Both cases pass.
  4. Save prompt v2 with text Reply: {{query}}. Run it on the same revision; both checks fail. Open the candidate from Run history, choose v1 as the reference run, and filter regressed cases.
  5. Promote v1 to production. Fetch it with the SDK below. Enable evaluation access on a separate CI key to try the gate.

Demo mode echoes the compiled prompt; it does not call a generation model or measure model quality. For a live use case, test a support-ticket classifier using JSON labels, references, and JSON-path checks.

Install and fetch a release

The Python package is included in this repository and is not published to PyPI. Create a project API key in Settings. Default keys read prompts; evaluation access is optional and must be enabled when creating the key.

# From the LLMForge repository root (Python 3.10+)
python -m pip install -e ./sdk/python

# Set environment variables in your shell / CI secret store:
# LLMFORGE_URL=http://localhost:8000/api/v1
# LLMFORGE_API_KEY=<project key from Settings>
from llmforge import LLMForge

with LLMForge() as forge:
    prompt = forge.get_prompt("Echo demo", label="production")
    print(prompt.compile(query="Hello"))
    # Fetch a pinned snapshot with version=1 instead of label=...

Default fetch selects production. Staging and explicit versions are supported. The SDK revalidates its memory cache with the server on every fetch. Missing releases, revoked keys, and network failures raise errors without silently selecting another version.

Gate changes in CI

Use a project key with evaluations:write. Pin a prompt version and dataset revision for reproducible results, or select staging to evaluate the current candidate. CI needs a reachable backend with a persistent process.

python -m llmforge evaluate \
  --prompt "Echo demo" --version 1 \
  --dataset Greetings --dataset-version 1 \
  --min-pass-rate 1 --output artifacts/evaluation.json

python -m llmforge check artifacts/evaluation.json

Exit 0: passed. Exit 1: quality regression or case errors. Exit 2: configuration, transport, timeout, or incomplete-run failure. Reports can be uploaded as CI artifacts; they include dataset inputs and outputs.

For live generation set LLMFORGE_PROVIDER_API_KEY and select --provider / --model. Live calls use your provider credits. The CLI attempts cancellation on timeout.

Copy docs/examples/evaluation-gate.yml from the repository for a complete GitHub Actions workflow. Set the backend URL as a repository variable and the evaluation key as a secret.

Evaluate in your own application

Use get_dataset to load fixed cases, run your application or model locally, and submit one output and check set per case with submit_evaluation. Include named numeric metrics if useful. External results appear in the same grid, explicitly marked as application-supplied checks. They are self-reported and do not trigger provider calls.

run_id = forge.submit_evaluation(
    prompt.id, dataset["revision"]["id"],
    model="your-application", results=results,
    metrics={"accuracy": 0.95},
)
result = forge.get_evaluation(run_id)
assert result.meets_threshold(0.9)

Each result includes case_index, output, and checks such as {"name": "correctness", "passed": true}, or an error. See the complete example in sdk/python/examples/submit_results.py.

Limits and troubleshooting

  • 401: check your project key and whether it was revoked. Clerk browser tokens and project API keys serve different endpoints.
  • 403: create a new key with evaluation access enabled. Existing read-only keys remain read-only.
  • 404: verify the project, exact name, release label and version. Promote a prompt before fetching production.
  • 422: check variable names, references, assertions, and case limits. Exact match requires a reference for every case.
  • Dataset revisions allow 100 cases / 1 MB. Live runs allow 50 cases, or 20 with judging. Two runs can be active per project.
  • Provider/judge calls have a 30-second limit and no automatic retry. Restarting the backend loses in-memory credentials. Interrupted evaluations preserve committed checkpoints and offer explicit resume with fresh keys; a provider response lost before saving may be charged again.
  • LLM judge scores are model opinions. Prefer deterministic checks for strict JSON or output contracts, and inspect judge reasons.

The packaged SDK and guide cover this personal project's current workflow. Reasoning benchmarks remain available in the separate Benchmarks area.