Evaluate a version for the promotion gate
Creates the experiment the lineage’s promotion gate requires for this version: the gate’s dataset snapshot, exactly its pinned evaluator versions and its trials, targeting the version (resolved to the fleet adapter or single-model deployment serving exactly those weights). When another version is in production, its latest matching experiment is the baseline, so the run carries its diff. Works whether or not the gate is enabled. Promote once it completes.
curl --request POST \
--url https://api.veri.studio/v1/models/{name}/versions/{version}/evaluate \
--header 'Authorization: Bearer <token>'import requests
url = "https://api.veri.studio/v1/models/{name}/versions/{version}/evaluate"
headers = {"Authorization": "Bearer <token>"}
response = requests.post(url, headers=headers)
print(response.text)const options = {method: 'POST', headers: {Authorization: 'Bearer <token>'}};
fetch('https://api.veri.studio/v1/models/{name}/versions/{version}/evaluate', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"object": "<string>",
"id": "<string>",
"name": "<string>",
"status": "<string>",
"dataset_id": "<string>",
"dataset_snapshot_id": "<string>",
"target": {},
"evaluators": [
{
"evaluator_id": "<string>",
"version": 123,
"name": "<string>",
"scope": "<string>",
"method": "<string>",
"judge_model": "<string>"
}
],
"bindings": {},
"generation": {
"temperature": 123,
"max_tokens": 123
},
"trials": 123,
"items_total": 123,
"items_done": 123,
"created_at": "2023-11-07T05:31:56Z",
"target_deployment_id": "<string>",
"target_model": "<string>",
"target_model_version_id": "<string>",
"target_served_model": "<string>",
"aggregates": {},
"passed": true,
"baseline_experiment_id": "<string>",
"diff": {},
"error": "<string>",
"created_by": "<string>",
"started_at": "2023-11-07T05:31:56Z",
"finished_at": "2023-11-07T05:31:56Z"
}Authorizations
API key with the vk_ prefix. Create one from the dashboard.
Response
Accepted (status queued); poll GET /v1/experiments/{id}
One experiment.
Always "experiment".
exp_...
queued | running | awaiting_local | completed | failed | canceled
The snapshot the run is pinned to (<dataset_id>@snap-N).
The target as requested.
The pinned evaluator versions (+ each judge's served model).
Show child attributes
Show child attributes
Sampling params for the target's calls. Omitted = the server default.
Show child attributes
Show child attributes
Items = rows x trials; done counts finished items (ok or error).
Resolved at create: the deployment called, the request model sent
(an adapter's versioned internal name for a model version on a
fleet), the model version, and the served model's identity.
Per evaluator at completion: n, ok, errors, skipped, pass_rate (all items in the denominator), mean, label_counts, critical_failures.
Every evaluator with a pass rule passed every item with no errors; null until completed, or when no evaluator has a pass rule.
The diff against the baseline, computed at completion.
curl --request POST \
--url https://api.veri.studio/v1/models/{name}/versions/{version}/evaluate \
--header 'Authorization: Bearer <token>'import requests
url = "https://api.veri.studio/v1/models/{name}/versions/{version}/evaluate"
headers = {"Authorization": "Bearer <token>"}
response = requests.post(url, headers=headers)
print(response.text)const options = {method: 'POST', headers: {Authorization: 'Bearer <token>'}};
fetch('https://api.veri.studio/v1/models/{name}/versions/{version}/evaluate', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"object": "<string>",
"id": "<string>",
"name": "<string>",
"status": "<string>",
"dataset_id": "<string>",
"dataset_snapshot_id": "<string>",
"target": {},
"evaluators": [
{
"evaluator_id": "<string>",
"version": 123,
"name": "<string>",
"scope": "<string>",
"method": "<string>",
"judge_model": "<string>"
}
],
"bindings": {},
"generation": {
"temperature": 123,
"max_tokens": 123
},
"trials": 123,
"items_total": 123,
"items_done": 123,
"created_at": "2023-11-07T05:31:56Z",
"target_deployment_id": "<string>",
"target_model": "<string>",
"target_model_version_id": "<string>",
"target_served_model": "<string>",
"aggregates": {},
"passed": true,
"baseline_experiment_id": "<string>",
"diff": {},
"error": "<string>",
"created_by": "<string>",
"started_at": "2023-11-07T05:31:56Z",
"finished_at": "2023-11-07T05:31:56Z"
}
