Update dataset rule
Set or clear the auto-snapshot rule: with auto_snapshot_rows = N, the platform cuts a snapshot once N new rows have landed since the newest snapshot (fires once per threshold crossing).
curl --request PATCH \
--url https://api.veri.studio/v1/datasets/{dataset_id} \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"auto_snapshot_rows": 123
}
'import requests
url = "https://api.veri.studio/v1/datasets/{dataset_id}"
payload = { "auto_snapshot_rows": 123 }
headers = {
"Authorization": "Bearer <token>",
"Content-Type": "application/json"
}
response = requests.patch(url, json=payload, headers=headers)
print(response.text)const options = {
method: 'PATCH',
headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},
body: JSON.stringify({auto_snapshot_rows: 123})
};
fetch('https://api.veri.studio/v1/datasets/{dataset_id}', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"object": "<string>",
"id": "<string>",
"name": "<string>",
"created_at": "2023-11-07T05:31:56Z",
"source_type": "<string>",
"source_uri": "<string>",
"huggingface_dataset": "<string>",
"num_rows": 123,
"format": "<string>",
"head_row": 123,
"auto_snapshot_rows": 123,
"latest_snapshot_id": "<string>",
"unsnapshotted_rows": 123,
"snapshot_count": 123
}Authorizations
API key with the vk_ prefix. Create one from the dashboard.
Path Parameters
Dataset ID or name
Body
null disables the auto-snapshot rule; N >= 1 enables it.
Response
The updated dataset
Row format. A stream's locked format (prompt | preference | completion |
chat | eval); an upload records its first row's shape, which adds
chat_no_prompt (chat rows without a prompt, still locked as chat), text,
sharegpt and alpaca (no lock). Job submit refuses GRPO on chat_no_prompt,
text, sharegpt and alpaca rows. Null on Hugging Face / volume sources, and
on a legacy dataset until its first append locks it.
Highest row_id appended so far (0 = empty stream). Includes supersede and tombstone entries, so it is the log length, not the live row count.
"cut a snapshot every N new rows" rule; null = off.
Newest snapshot id ("@snap-N"); populated by GET /v1/datasets/{id}.
Rows appended since the newest snapshot; populated by GET /v1/datasets/{id}.
curl --request PATCH \
--url https://api.veri.studio/v1/datasets/{dataset_id} \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"auto_snapshot_rows": 123
}
'import requests
url = "https://api.veri.studio/v1/datasets/{dataset_id}"
payload = { "auto_snapshot_rows": 123 }
headers = {
"Authorization": "Bearer <token>",
"Content-Type": "application/json"
}
response = requests.patch(url, json=payload, headers=headers)
print(response.text)const options = {
method: 'PATCH',
headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},
body: JSON.stringify({auto_snapshot_rows: 123})
};
fetch('https://api.veri.studio/v1/datasets/{dataset_id}', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"object": "<string>",
"id": "<string>",
"name": "<string>",
"created_at": "2023-11-07T05:31:56Z",
"source_type": "<string>",
"source_uri": "<string>",
"huggingface_dataset": "<string>",
"num_rows": 123,
"format": "<string>",
"head_row": 123,
"auto_snapshot_rows": 123,
"latest_snapshot_id": "<string>",
"unsnapshotted_rows": 123,
"snapshot_count": 123
}
