API documentation
Use one ruble-denominated account for OpenAI-compatible inference, safe account automation and agent tools.
https://zapros.ai/api/v1
Inference endpoints accept zpr_ keys. Management REST and MCP accept a manually provisioned zpm_ token or a short-lived OAuth zpo_ issued for their own exact resource. MCP and Management OAuth use separate grants and tokens; never interchange their audiences.
Quickstart
- 1. Create an account. Register at zapros.ai/en/register.
- 2. Create an inference key. Open API keys in the dashboard and store the one-time
zpr_...secret securely. - 3. Top up and send a request. Requests are charged against your ruble balance after actual provider usage.
curl https://zapros.ai/api/v1/chat/completions \
-H "Authorization: Bearer zpr_your_key" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-5-nano",
"messages": [{"role":"user","content":"Hello!"}],
"max_tokens": 512
}'zpr_ and zpm_secrets on trusted servers. Newly created secrets are shown once and cannot be recovered.Authentication
Inference endpoints use Authorization: Bearer zpr_.... Management REST and MCP accept separately scoped, manually provisioned zpm_... tokens. Each surface also accepts short-lived zpo_... access tokens issued for its exact resource. MCP and Management OAuth grants are separate and their tokens are not interchangeable.
Chat Completions
POST /api/v1/chat/completions supports the familiar OpenAI request format, function tools and SSE streaming. Exact parameter support depends on the selected model. Always set a finite output limit for generative requests.
from openai import OpenAI
client = OpenAI(
api_key="zpr_your_key",
base_url="https://zapros.ai/api/v1",
)
response = client.chat.completions.create(
model="openai/gpt-5-nano",
messages=[{"role": "user", "content": "Hello!"}],
max_tokens=512,
)
print(response.choices[0].message.content)Responses API
POST /api/v1/responses supports reasoning, structured outputs, client function tools, image/PDF input and streaming events. Use store: false and pass complete history in input;previous_response_id is not supported.
const response = await client.responses.create({
model: "openai/gpt-5-nano",
input: "Explain quantum entanglement briefly",
max_output_tokens: 1024,
reasoning: { effort: "medium" },
store: false,
});
console.log(response.output_text);Bounded openrouter:web_search is available with exa, parallelor perplexity. Set finitemax_uses, max_resultsand max_total_results. Legacy web plugins, X search, audio/video input and output-image tools are rejected.
Embeddings
POST /api/v1/embeddings accepts one string or an array of strings. Safe text-only requests are billed by input tokens; media, token-ID arrays, dimensionsand input_type are currently rejected.
const result = await client.embeddings.create({
model: "openai/text-embedding-3-small",
input: ["First document", "Second document"],
});
console.log(result.data[0].embedding);Video jobs
Start an asynchronous job with POST /api/v1/videos, then poll GET /api/v1/videos/{id} using a key from the same account. Video is billed per second and failed jobs are refunded automatically.
curl https://zapros.ai/api/v1/videos \
-H "Authorization: Bearer zpr_your_key" \
-H "Content-Type: application/json" \
-d '{
"model": "google/veo-3.1-fast",
"prompt": "A calm aerial shot over a forest",
"duration": 8
}'Streaming
Set stream: true to receive incremental Server-Sent Events. Chat Completions emits choices[].delta; Responses emits typed events such as response.output_text.delta.
const stream = await client.chat.completions.create({
model: "openai/gpt-5-nano",
messages: [{ role: "user", content: "Hello!" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}The final usage event is metered automatically. If a connection is interrupted, retry only according to the endpoint semantics; never concatenate a blind duplicate request.
Models & pricing
Use the live catalogs instead of hard-coding a stale model list. Catalogs fail closed if upstream pricing data becomes stale.
# Chat Completions and Responses
curl https://zapros.ai/api/v1/models
# Embeddings
curl https://zapros.ai/api/v1/embeddings/models
# Video
curl https://zapros.ai/api/v1/videosPricing fields explicitly identify token rates and fixed request, image and web-search fees. A search engine can add its own fee. Browse the same live data on Models & pricing.
Limits & billing
- • Rate limit: 200 requests per minute per inference key. Respect HTTP 429 and
Retry-After. - • Balance and budgets: HTTP 402 means the balance or a configured key budget is insufficient for the bounded reservation.
- • Per-key controls: configure a monthly ruble budget, a per-request ruble cap and a model allowlist. In-flight reservations count toward the budget.
- • Settlement: final charges use provider usage where available; unused holds are released and definite pre-processing failures are refunded.
Management API
Management automation can use the separate zpm_namespace. Create and revoke those tokens from Agents & MCP. Public OAuth clients instead request the exact resource https://zapros.ai/api/v1/management. That grant is separate from MCP OAuth, and a token for one resource fails at the other. Neither credential can create management tokens, access payments or top up a balance.
GET https://zapros.ai/.well-known/oauth-protected-resource/api/v1/management
GET https://zapros.ai/.well-known/oauth-authorization-server
POST https://zapros.ai/oauth/register
GET https://zapros.ai/oauth/authorize?resource=https%3A%2F%2Fzapros.ai%2Fapi%2Fv1%2Fmanagement
POST https://zapros.ai/oauth/token| Scope | Access |
|---|---|
| account:read | Account identity and RUB balance |
| models:read | Catalogs and cost estimation |
| keys:read | Inference-key metadata and limits |
| keys:create | Create bounded inference keys |
| keys:update | Update inference keys |
| keys:revoke | Permanently revoke inference keys |
| usage:read | Usage totals and recent requests |
| compute:read | Current and future CPU/GPU resources |
curl "https://zapros.ai/api/v1/management/models?kind=video&limit=20" \
-H "Authorization: Bearer zpm_your_management_token"
curl https://zapros.ai/api/v1/management/cost-estimates \
-X POST \
-H "Authorization: Bearer zpm_your_management_token" \
-H "Content-Type: application/json" \
-d '{"endpoint":"chat","model":"openai/gpt-5-nano","input_tokens":2000,"output_tokens":500}'Both endpoints require models:read. An estimate uses current catalog pricing and never calls a provider, reserves balance, or charges the account; an accepted request's actual settled cost may differ.
curl "https://zapros.ai/api/v1/management/projects?limit=50" \
-H "Authorization: Bearer zpm_your_management_token"
# If the list is empty, safely create the personal default project:
curl https://zapros.ai/api/v1/management/projects/default \
-X POST \
-H "Authorization: Bearer zpm_your_management_token"GET /projects works with any key-management scope. The idempotent bootstrap requires keys:create. Every inference-key operation then requires an explicit project_id; keys, limits, and usage remain isolated by project.
curl https://zapros.ai/api/v1/management/keys \
-X POST \
-H "Authorization: Bearer zpm_your_management_token" \
-H "Idempotency-Key: agent-key-2026-10-04" \
-H "Content-Type: application/json" \
-d '{
"project_id": "project-id-from-the-previous-response",
"name": "Research agent",
"monthlyBudgetRub": 500,
"maxRequestRub": 10,
"allowedModels": ["openai/gpt-5-nano"]
}'A trusted direct zpm_ caller receives the new zpr_ secret once. A Management OAuth client never receives a raw zpr_; it receives only a 15-minute dashboard claim for the account owner. A stable 8–128 character Idempotency-Key prevents duplicate creation for 24 hours, but it cannot restore a secret after its dashboard claim was acknowledged or expired. Such a terminal replay returns claim metadata without the secret. Management mutations are audited.
GET /api/v1/management/account
GET /api/v1/management/models?kind=all&limit=50
POST /api/v1/management/cost-estimates
GET /api/v1/management/projects?limit=50
POST /api/v1/management/projects/default
GET /api/v1/management/keys?project_id={project_id}
POST /api/v1/management/keys
PATCH /api/v1/management/keys/{id}
DELETE /api/v1/management/keys/{id}?project_id={project_id}
GET /api/v1/management/usage?days=30&limit=50
GET /api/v1/management/compute/offers
GET /api/v1/management/compute/instancesMCP for agents
Zapros exposes a stateless, POST-only JSON-RPC MCP server at https://zapros.ai/mcp. It supports MCP 2026-07-28 and safe compatibility with 2025-11-25 and 2025-06-18. There is no GET/SSE session stream. The preferred connection method is to give a compatible client only the MCP URL and complete Agent Connect OAuth. Ready-to-use Claude, Cursor and Codex instructions are available on the Connect an agent page. A manually provisioned zpm_ bearer remains supported for explicit configuration.
Agent Connect OAuth
The MCP authentication challenge points to RFC 9728 protected-resource metadata, which links to RFC 8414 authorization-server metadata. Public clients use authorization code with S256 PKCE, an exact registered redirect URI, and resource=https://zapros.ai/mcp. Client secrets are neither required nor accepted. Claude and Codex can use a Client ID Metadata Document, while Cursor uses bounded dynamic public-client registration.
GET https://zapros.ai/.well-known/oauth-protected-resource/mcp
GET https://zapros.ai/.well-known/oauth-authorization-server
POST https://zapros.ai/oauth/register
GET https://zapros.ai/oauth/authorize
POST https://zapros.ai/oauth/tokenA zpo_ access token normally lasts 15 minutes. Refresh tokens rotate on every exchange; reusing an old refresh token revokes that MCP grant. This token works only at MCP. Management REST requires a separate OAuth grant with resource=https://zapros.ai/api/v1/management; audiences are not interchangeable.
Manual connection
Create a scoped zpm_ token in the dashboard and configure it as the MCP bearer header. This account-scoped credential is intended for trusted automation and is distinct from a consent-bounded OAuth grant.
curl https://zapros.ai/mcp \
-X POST \
-H "Authorization: Bearer zpm_your_management_token" \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-H "MCP-Protocol-Version: 2026-07-28" \
-H "Mcp-Method: server/discover" \
-d '{
"jsonrpc": "2.0",
"id": 1,
"method": "server/discover",
"params": {
"_meta": {
"io.modelcontextprotocol/protocolVersion": "2026-07-28",
"io.modelcontextprotocol/clientInfo": {
"name": "my-agent",
"version": "1.0.0"
},
"io.modelcontextprotocol/clientCapabilities": {}
}
}
}'Modern methods are server/discover,ping, tools/listand tools/call. For a tool call, setMcp-Method: tools/call and anMcp-Name header matchingparams.name. The available tool list is filtered by token scope. Call list_projectsfirst (or the idempotent ensure_default_projectwith keys:create), then pass the returned project object's id as project_id to every key tool.
get_account
Account identity and balance
list_projects
Discover key-owning projects
ensure_default_project
Idempotent personal-project bootstrap
list_models
Paginated current model catalog
estimate_cost
Bounded RUB cost estimate
list_api_keys
Inference-key metadata
create_api_key
24-hour retry-safe key creation
update_api_key
Name and limit updates
revoke_api_key
Permanent key revocation
get_usage
Usage aggregates and recent requests
list_compute_offers
CPU/GPU offer discovery
list_compute_instances
Account compute inventory
keys:create, keys:update, and keys:revoke scopes. MCP never puts a raw new zpr_ secret in model-facing output; it returns an owner-authenticated dashboard claim. The owner can safely retry Reveal for 15 minutes and explicitly confirms after saving the key; without confirmation, expiry disables the key. Grant only the operations the client needs.Compute discovery
compute:read already provides stable discovery endpoints and MCP tools for future CPU/GPU hosting. Until hosting is enabled, an empty offers or instances list is a valid successful response—not an outage. Creation, shutdown and billing mutations are not exposed yet.
Errors
APIs return machine-readable JSON errors and a request ID. Log the request ID, but never log bearer tokens or full prompts.
| HTTP | Meaning | Action |
|---|---|---|
| 400 | Invalid or unsupported input | Fix the request; do not retry unchanged |
| 401 | Missing or invalid secret | Use the correct zpr_/zpm_/zpo_ namespace |
| 402 | Balance or key limit exceeded | Top up or adjust the relevant limit |
| 403 | Missing management scope | Issue a least-privilege token with that scope |
| 429 | Rate limited | Back off and respect Retry-After |
| 5xx | Transient service/provider error | Retry boundedly with idempotency where supported |
Code examples
Zapros works with OpenAI-compatible SDKs or ordinary HTTP clients. Keep the base URL and bearer-key namespace explicit.
from openai import OpenAI
client = OpenAI(
api_key="zpr_your_key",
base_url="https://zapros.ai/api/v1",
)
response = client.chat.completions.create(
model="openai/gpt-5-nano",
messages=[{"role": "user", "content": "Hello!"}],
max_tokens=512,
)
print(response.choices[0].message.content)const response = await fetch(
"https://zapros.ai/api/v1/chat/completions",
{
method: "POST",
headers: {
Authorization: "Bearer zpr_your_key",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "openai/gpt-5-nano",
messages: [{ role: "user", content: "Hello!" }],
max_tokens: 512,
}),
},
);
const data = await response.json();
console.log(data.choices[0].message.content);Machine-readable references
Need help? Contact Telegram or support@zapros.ai.
Get an API key
zapros.ai