Appearance
Leaderboards API
The resource API lets backend services manage leaderboards, models, runs, and schedules.
bash
export API_BASE="https://dr-gero-frontend-99142474693.europe-west1.run.app"
export DRGERO_TOKEN="drgero_REPLACE_WITH_TOKEN_FROM_SETTINGS"List leaderboards
bash
curl -sS "$API_BASE/api/leaderboards?limit=50&offset=0" \
-H "Authorization: Bearer $DRGERO_TOKEN" | jqRequires leaderboards:read.
Create a GET-dataset leaderboard
bash
curl -sS -X POST "$API_BASE/api/leaderboards" \
-H "Authorization: Bearer $DRGERO_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Support QA",
"system_prompt": "Answer the support question clearly.\n\nQuestion:\n{input}",
"dataset_type": "GET",
"dataset_url": "https://huggingface.co/datasets/acme/support-evals/resolve/main/eval.jsonl",
"eval_type": "judge",
"judge_auto_decide": false,
"judge_provider": "OpenRouter",
"judge_model": "openai/gpt-5.2"
}' | jqRequires leaderboards:write.
Required fields:
| Field | Required | Notes |
|---|---|---|
name | Yes | Leaderboard/challenge name. |
system_prompt or model_prompt | Yes | Task prompt. CamelCase aliases are accepted in several places. |
dataset_type | No | GET by default. Use PUSH for webhook datasets. |
dataset_url | Yes for GET | Must be a Hugging Face JSONL URL. |
eval_type | No | exact, judge, or human. |
judge_auto_decide | No | Defaults to true. Set to false to provide a manual judge. |
judge_provider, judge_model | When judge_auto_decide is false | Manual judge configuration. |
Create a PUSH-dataset leaderboard
bash
curl -sS -X POST "$API_BASE/api/leaderboards" \
-H "Authorization: Bearer $DRGERO_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Production Support Feedback",
"system_prompt": "Answer the user request.\n\nInput:\n{input}",
"dataset_type": "PUSH",
"eval_type": "judge",
"judge_auto_decide": false,
"judge_provider": "OpenRouter",
"judge_model": "openai/gpt-5.2",
"auto_limit_size": true,
"max_samples_to_gather": 1000,
"daily_event_limit": 5000,
"monthly_event_limit": 50000,
"consolidate_every_events": 500,
"consolidate_every_hours": 24,
"dedupe": true
}' | jqGet leaderboard detail
bash
curl -sS "$API_BASE/api/leaderboards/$LEADERBOARD_ID" \
-H "Authorization: Bearer $DRGERO_TOKEN" | jqThe response includes leaderboard metadata, challenge, candidate models, and recent runs.
Update a leaderboard
bash
curl -sS -X PATCH "$API_BASE/api/leaderboards/$LEADERBOARD_ID" \
-H "Authorization: Bearer $DRGERO_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"description": "Production support leaderboard",
"chosen_model_strategy": "best_balanced"
}' | jqUpdatable fields include:
name,description,statusdataset_type,dataset_url,dataset_metadata,dataset_auto_labelseval_type,eval_metadatacategory,constraints,model_prompt,system_promptchosen_leaderboard_model_id,chosen_model_strategyschedule
Schedule updates may require a paid entitlement.
chosen_model_strategy accepts:
best_balanced(default): use Gero-0 balanced routing when an eligible mix is available, otherwise fall back to the ranking winner.ranking_winner: always follow the latest ranking winner.manual: usechosen_leaderboard_model_id.
Delete a leaderboard
bash
curl -sS -X DELETE "$API_BASE/api/leaderboards/$LEADERBOARD_ID" \
-H "Authorization: Bearer $DRGERO_TOKEN" | jqFree-plan workspaces may be prevented from deleting leaderboards.
Add a candidate model
bash
curl -sS -X POST "$API_BASE/api/leaderboards/$LEADERBOARD_ID/models" \
-H "Authorization: Bearer $DRGERO_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "GPT OSS 120B via OpenRouter",
"platform": "OpenRouter",
"model_url": "openai/gpt-oss-120b"
}' | jqManual candidate fields:
| Field | Required | Notes |
|---|---|---|
name or model_name | Yes | Display name. |
platform | No | OpenRouter, Custom, HuggingFace, or Dr.Gero. Defaults to Dr.Gero. |
model_url or url | Required except Dr.Gero | OpenRouter model ID or endpoint URL. |
model_id | Required for Dr.Gero | ID of a Dr.Gero model. |
token | No | Model-specific token if not using workspace integration. Returned redacted. |
auth_type | No | bearer, x-api-key, x-dr.gero-api-key, authorization, or custom-header. |
auth_header_name | For custom header | Header name for custom auth. |
Adding models may require a paid entitlement.
Candidate models cannot be added while the leaderboard is running. When the candidate is a Dr.Gero model, the server verifies that it has a successful fine tune, an active deployment, and can answer a sampled leaderboard input. A model that fails this check is stored as pre_production and is excluded from leaderboard inference. Posting the same Dr.Gero model again re-runs the check and returns the updated candidate with rechecked: true.
Auto-select candidate models
Auto-select is a UI-session endpoint, because it uses the signed-in workspace context and OpenRouter integration. It is useful for browser/admin automation rather than server-to-server API-token automation.
bash
curl -sS -X POST "$API_BASE/api/leaderboards/$LEADERBOARD_ID/models/auto-select" \
-H "Authorization: Bearer $SUPABASE_ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"number_of_models": 5,
"limit_cost": true,
"input_price_per_m_tokens": 0.5,
"output_price_per_m_tokens": 1,
"limit_latency": true,
"latency_p95_seconds": 1,
"latency_p99_seconds": 3,
"only_open_source": false
}' | jqBody fields accept camelCase aliases such as numberOfModels, limitCost, inputPricePerMTokens, latencyP95Seconds, and onlyOpenSource.
List and remove candidate models
bash
curl -sS "$API_BASE/api/leaderboards/$LEADERBOARD_ID/models" \
-H "Authorization: Bearer $DRGERO_TOKEN" | jq
curl -sS -X DELETE "$API_BASE/api/leaderboards/$LEADERBOARD_ID/models/$LEADERBOARD_MODEL_ID" \
-H "Authorization: Bearer $DRGERO_TOKEN" | jqRun a leaderboard
bash
curl -sS -X POST "$API_BASE/api/leaderboards/$LEADERBOARD_ID/run" \
-H "Authorization: Bearer $DRGERO_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model_ids": ["LEADERBOARD_MODEL_ID_1", "LEADERBOARD_MODEL_ID_2"],
"run_source": "manual",
"max_examples": 1000,
"example_selection_algorithm": "LAST_N"
}' | jqRequires leaderboards:run.
Use /api/leaderboards/{leaderboard_id}/improve-dataset to force a dataset-improvement run.
max_examples is an optional positive integer. example_selection_algorithm can be FIRST_N, LAST_N, RANDOM_SHUFFLE, or WEIGHTED_SHUFFLE. The same fields are accepted by the dataset-improvement route.
Dataset enhancement
Set improve_dataset: true when creating or updating a leaderboard to persist the feature, or include it in a run request to enable enhancement for that run. For explicit persistent control during create/update, use data_enhancement:
json
{
"data_enhancement": {
"enabled": true,
"strict": false,
"use_existing_scores": true,
"max_judgements": 200,
"max_repairs": 200,
"judge": {"enabled": false},
"repair": {"enabled": true},
"sft": {"enabled": true},
"synthetic": {"enabled": false, "per_repaired_row": 0, "max_rows": 0},
"publish": {"mode": "artifact_only"}
}
}The normalized configuration is stored under dataset_metadata.data_enhancement.
Run artifacts
List retained artifacts produced by leaderboard runs:
bash
curl -sS "$API_BASE/api/leaderboards/$LEADERBOARD_ID/artifacts" \
-H "Authorization: Bearer $DRGERO_TOKEN" | jqRequires leaderboards:read. Artifact metadata indicates whether the file is still downloadable. Request a short-lived download URL with:
bash
curl -sS \
"$API_BASE/api/leaderboards/$LEADERBOARD_ID/artifacts/$ARTIFACT_ID/download" \
-H "Authorization: Bearer $DRGERO_TOKEN" | jqRun and inference webhooks
Leaderboards support independent webhooks for run events and inference events:
bash
curl -sS -X PATCH \
"$API_BASE/api/leaderboards/$LEADERBOARD_ID/run/webhook" \
-H "Authorization: Bearer $DRGERO_TOKEN" \
-H "Content-Type: application/json" \
-d '{"webhook_url":"https://example.com/dr-gero/runs","enabled":true}' | jq
curl -sS -X PATCH \
"$API_BASE/api/leaderboards/$LEADERBOARD_ID/inference/webhook" \
-H "Authorization: Bearer $DRGERO_TOKEN" \
-H "Content-Type: application/json" \
-d '{"webhook_url":"https://example.com/dr-gero/inference","enabled":true}' | jqPATCH, PUT, and POST are accepted and require leaderboards:write. Webhook URLs must be public HTTPS URLs without embedded credentials. Send an empty or null webhook_url to disable delivery.
Run delivery emits leaderboard.run.trace events followed by a terminal leaderboard.run.completed, leaderboard.run.failed, or leaderboard.run.canceled event. Inference delivery emits leaderboard.inference. Deliveries are best-effort.
Schedule JSON
Schedules are stored on the leaderboard with a JSON structure like:
json
{
"version": 2,
"triggers": {
"model_version": { "enabled": true, "cadence": "WEEKLY" },
"new_data": { "enabled": true, "check_cadence": "DAILY", "every_new_events": 500 },
"cron": { "enabled": true, "preset": "CUSTOM", "expression": "0 6 * * 1" }
},
"dataset": {
"mode": "LIMIT",
"limit": { "auto": false, "rows": 5000, "algorithm": "LAST_N" }
}
}Save it with:
bash
curl -sS -X PATCH "$API_BASE/api/leaderboards/$LEADERBOARD_ID" \
-H "Authorization: Bearer $DRGERO_TOKEN" \
-H "Content-Type: application/json" \
-d @schedule.json | jqA schedule cannot be supplied during leaderboard creation. Create the leaderboard, complete one successful run, and then save the schedule with PATCH.
Human evaluation differences
Create Human Eval through the UI workflow to start without dataset rows. The resource creation examples above retain their GET/PUSH validation; eval_type: human alone does not remove the GET URL requirement on that endpoint.
Human cases use the session-authenticated Human evaluation API. Auto-select works without data for this mode. Automatic runs, dataset-improvement runs and schedules are blocked. Serving follows the common-case winner rather than Best Balanced or a manually pinned model. Once human cases exist, changing the evaluation method is blocked; create another leaderboard instead.