API documentation

Use one ruble-denominated account for OpenAI-compatible inference, safe account automation and agent tools.

Inference base URL

https://zapros.ai/api/v1

Inference endpoints accept zpr_ keys. Management REST and MCP accept a manually provisioned zpm_ token or a short-lived OAuth zpo_ issued for their own exact resource. MCP and Management OAuth use separate grants and tokens; never interchange their audiences.

Quickstart

  1. 1. Create an account. Register at zapros.ai/en/register.
  2. 2. Create an inference key. Open API keys in the dashboard and store the one-time zpr_... secret securely.
  3. 3. Top up and send a request. Requests are charged against your ruble balance after actual provider usage.
curl
curl https://zapros.ai/api/v1/chat/completions \
  -H "Authorization: Bearer zpr_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5-nano",
    "messages": [{"role":"user","content":"Hello!"}],
    "max_tokens": 512
  }'
Keep zpr_ and zpm_secrets on trusted servers. Newly created secrets are shown once and cannot be recovered.

Authentication

Inference endpoints use Authorization: Bearer zpr_.... Management REST and MCP accept separately scoped, manually provisioned zpm_... tokens. Each surface also accepts short-lived zpo_... access tokens issued for its exact resource. MCP and Management OAuth grants are separate and their tokens are not interchangeable.

Chat Completions

POST /api/v1/chat/completions supports the familiar OpenAI request format, function tools and SSE streaming. Exact parameter support depends on the selected model. Always set a finite output limit for generative requests.

Python
from openai import OpenAI

client = OpenAI(
    api_key="zpr_your_key",
    base_url="https://zapros.ai/api/v1",
)

response = client.chat.completions.create(
    model="openai/gpt-5-nano",
    messages=[{"role": "user", "content": "Hello!"}],
    max_tokens=512,
)
print(response.choices[0].message.content)

Responses API

POST /api/v1/responses supports reasoning, structured outputs, client function tools, image/PDF input and streaming events. Use store: false and pass complete history in input;previous_response_id is not supported.

JavaScript
const response = await client.responses.create({
  model: "openai/gpt-5-nano",
  input: "Explain quantum entanglement briefly",
  max_output_tokens: 1024,
  reasoning: { effort: "medium" },
  store: false,
});

console.log(response.output_text);

Bounded openrouter:web_search is available with exa, parallelor perplexity. Set finitemax_uses, max_resultsand max_total_results. Legacy web plugins, X search, audio/video input and output-image tools are rejected.

Embeddings

POST /api/v1/embeddings accepts one string or an array of strings. Safe text-only requests are billed by input tokens; media, token-ID arrays, dimensionsand input_type are currently rejected.

JavaScript
const result = await client.embeddings.create({
  model: "openai/text-embedding-3-small",
  input: ["First document", "Second document"],
});

console.log(result.data[0].embedding);

Video jobs

Start an asynchronous job with POST /api/v1/videos, then poll GET /api/v1/videos/{id} using a key from the same account. Video is billed per second and failed jobs are refunded automatically.

curl
curl https://zapros.ai/api/v1/videos \
  -H "Authorization: Bearer zpr_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/veo-3.1-fast",
    "prompt": "A calm aerial shot over a forest",
    "duration": 8
  }'

Streaming

Set stream: true to receive incremental Server-Sent Events. Chat Completions emits choices[].delta; Responses emits typed events such as response.output_text.delta.

JavaScript
const stream = await client.chat.completions.create({
  model: "openai/gpt-5-nano",
  messages: [{ role: "user", content: "Hello!" }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

The final usage event is metered automatically. If a connection is interrupted, retry only according to the endpoint semantics; never concatenate a blind duplicate request.

Models & pricing

Use the live catalogs instead of hard-coding a stale model list. Catalogs fail closed if upstream pricing data becomes stale.

Catalogs
# Chat Completions and Responses
curl https://zapros.ai/api/v1/models

# Embeddings
curl https://zapros.ai/api/v1/embeddings/models

# Video
curl https://zapros.ai/api/v1/videos

Pricing fields explicitly identify token rates and fixed request, image and web-search fees. A search engine can add its own fee. Browse the same live data on Models & pricing.

Limits & billing

  • • Rate limit: 200 requests per minute per inference key. Respect HTTP 429 andRetry-After.
  • • Balance and budgets: HTTP 402 means the balance or a configured key budget is insufficient for the bounded reservation.
  • • Per-key controls: configure a monthly ruble budget, a per-request ruble cap and a model allowlist. In-flight reservations count toward the budget.
  • • Settlement: final charges use provider usage where available; unused holds are released and definite pre-processing failures are refunded.

Management API

Management automation can use the separate zpm_namespace. Create and revoke those tokens from Agents & MCP. Public OAuth clients instead request the exact resource https://zapros.ai/api/v1/management. That grant is separate from MCP OAuth, and a token for one resource fails at the other. Neither credential can create management tokens, access payments or top up a balance.

Management OAuth discovery
GET https://zapros.ai/.well-known/oauth-protected-resource/api/v1/management
GET https://zapros.ai/.well-known/oauth-authorization-server
POST https://zapros.ai/oauth/register
GET https://zapros.ai/oauth/authorize?resource=https%3A%2F%2Fzapros.ai%2Fapi%2Fv1%2Fmanagement
POST https://zapros.ai/oauth/token
ScopeAccess
account:readAccount identity and RUB balance
models:readCatalogs and cost estimation
keys:readInference-key metadata and limits
keys:createCreate bounded inference keys
keys:updateUpdate inference keys
keys:revokePermanently revoke inference keys
usage:readUsage totals and recent requests
compute:readCurrent and future CPU/GPU resources
Catalog and cost estimate
curl "https://zapros.ai/api/v1/management/models?kind=video&limit=20" \
  -H "Authorization: Bearer zpm_your_management_token"

curl https://zapros.ai/api/v1/management/cost-estimates \
  -X POST \
  -H "Authorization: Bearer zpm_your_management_token" \
  -H "Content-Type: application/json" \
  -d '{"endpoint":"chat","model":"openai/gpt-5-nano","input_tokens":2000,"output_tokens":500}'

Both endpoints require models:read. An estimate uses current catalog pricing and never calls a provider, reserves balance, or charges the account; an accepted request's actual settled cost may differ.

Discover or bootstrap a project
curl "https://zapros.ai/api/v1/management/projects?limit=50" \
  -H "Authorization: Bearer zpm_your_management_token"

# If the list is empty, safely create the personal default project:
curl https://zapros.ai/api/v1/management/projects/default \
  -X POST \
  -H "Authorization: Bearer zpm_your_management_token"

GET /projects works with any key-management scope. The idempotent bootstrap requires keys:create. Every inference-key operation then requires an explicit project_id; keys, limits, and usage remain isolated by project.

Create a constrained inference key
curl https://zapros.ai/api/v1/management/keys \
  -X POST \
  -H "Authorization: Bearer zpm_your_management_token" \
  -H "Idempotency-Key: agent-key-2026-10-04" \
  -H "Content-Type: application/json" \
  -d '{
    "project_id": "project-id-from-the-previous-response",
    "name": "Research agent",
    "monthlyBudgetRub": 500,
    "maxRequestRub": 10,
    "allowedModels": ["openai/gpt-5-nano"]
  }'

A trusted direct zpm_ caller receives the new zpr_ secret once. A Management OAuth client never receives a raw zpr_; it receives only a 15-minute dashboard claim for the account owner. A stable 8–128 character Idempotency-Key prevents duplicate creation for 24 hours, but it cannot restore a secret after its dashboard claim was acknowledged or expired. Such a terminal replay returns claim metadata without the secret. Management mutations are audited.

Management routes
GET    /api/v1/management/account
GET    /api/v1/management/models?kind=all&limit=50
POST   /api/v1/management/cost-estimates
GET    /api/v1/management/projects?limit=50
POST   /api/v1/management/projects/default
GET    /api/v1/management/keys?project_id={project_id}
POST   /api/v1/management/keys
PATCH  /api/v1/management/keys/{id}
DELETE /api/v1/management/keys/{id}?project_id={project_id}
GET    /api/v1/management/usage?days=30&limit=50
GET    /api/v1/management/compute/offers
GET    /api/v1/management/compute/instances

MCP for agents

Zapros exposes a stateless, POST-only JSON-RPC MCP server at https://zapros.ai/mcp. It supports MCP 2026-07-28 and safe compatibility with 2025-11-25 and 2025-06-18. There is no GET/SSE session stream. The preferred connection method is to give a compatible client only the MCP URL and complete Agent Connect OAuth. Ready-to-use Claude, Cursor and Codex instructions are available on the Connect an agent page. A manually provisioned zpm_ bearer remains supported for explicit configuration.

Agent Connect OAuth

The MCP authentication challenge points to RFC 9728 protected-resource metadata, which links to RFC 8414 authorization-server metadata. Public clients use authorization code with S256 PKCE, an exact registered redirect URI, and resource=https://zapros.ai/mcp. Client secrets are neither required nor accepted. Claude and Codex can use a Client ID Metadata Document, while Cursor uses bounded dynamic public-client registration.

OAuth discovery
GET https://zapros.ai/.well-known/oauth-protected-resource/mcp
GET https://zapros.ai/.well-known/oauth-authorization-server
POST https://zapros.ai/oauth/register
GET https://zapros.ai/oauth/authorize
POST https://zapros.ai/oauth/token

A zpo_ access token normally lasts 15 minutes. Refresh tokens rotate on every exchange; reusing an old refresh token revokes that MCP grant. This token works only at MCP. Management REST requires a separate OAuth grant with resource=https://zapros.ai/api/v1/management; audiences are not interchangeable.

Consent caps the aggregate monthly budget across every active key created by the connection, per-request spend, allowed models, maximum active keys, and grant lifetime. A connection can see, update, and revoke only keys it created. Users can revoke a connected app under Agents & MCP, immediately invalidating its tokens and disabling its active keys.

Manual connection

Create a scoped zpm_ token in the dashboard and configure it as the MCP bearer header. This account-scoped credential is intended for trusted automation and is distinct from a consent-bounded OAuth grant.

MCP 2026 discovery
curl https://zapros.ai/mcp \
  -X POST \
  -H "Authorization: Bearer zpm_your_management_token" \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -H "MCP-Protocol-Version: 2026-07-28" \
  -H "Mcp-Method: server/discover" \
  -d '{
    "jsonrpc": "2.0",
    "id": 1,
    "method": "server/discover",
    "params": {
      "_meta": {
        "io.modelcontextprotocol/protocolVersion": "2026-07-28",
        "io.modelcontextprotocol/clientInfo": {
          "name": "my-agent",
          "version": "1.0.0"
        },
        "io.modelcontextprotocol/clientCapabilities": {}
      }
    }
  }'

Modern methods are server/discover,ping, tools/listand tools/call. For a tool call, setMcp-Method: tools/call and anMcp-Name header matchingparams.name. The available tool list is filtered by token scope. Call list_projectsfirst (or the idempotent ensure_default_projectwith keys:create), then pass the returned project object's id as project_id to every key tool.

get_account

Account identity and balance

list_projects

Discover key-owning projects

ensure_default_project

Idempotent personal-project bootstrap

list_models

Paginated current model catalog

estimate_cost

Bounded RUB cost estimate

list_api_keys

Inference-key metadata

create_api_key

24-hour retry-safe key creation

update_api_key

Name and limit updates

revoke_api_key

Permanent key revocation

get_usage

Usage aggregates and recent requests

list_compute_offers

CPU/GPU offer discovery

list_compute_instances

Account compute inventory

MCP deliberately excludes payments, top-ups, management-token creation and compute mutations. Key mutations use separate keys:create, keys:update, and keys:revoke scopes. MCP never puts a raw new zpr_ secret in model-facing output; it returns an owner-authenticated dashboard claim. The owner can safely retry Reveal for 15 minutes and explicitly confirms after saving the key; without confirmation, expiry disables the key. Grant only the operations the client needs.

Compute discovery

compute:read already provides stable discovery endpoints and MCP tools for future CPU/GPU hosting. Until hosting is enabled, an empty offers or instances list is a valid successful response—not an outage. Creation, shutdown and billing mutations are not exposed yet.

Errors

APIs return machine-readable JSON errors and a request ID. Log the request ID, but never log bearer tokens or full prompts.

HTTPMeaningAction
400Invalid or unsupported inputFix the request; do not retry unchanged
401Missing or invalid secretUse the correct zpr_/zpm_/zpo_ namespace
402Balance or key limit exceededTop up or adjust the relevant limit
403Missing management scopeIssue a least-privilege token with that scope
429Rate limitedBack off and respect Retry-After
5xxTransient service/provider errorRetry boundedly with idempotency where supported

Code examples

Zapros works with OpenAI-compatible SDKs or ordinary HTTP clients. Keep the base URL and bearer-key namespace explicit.

Python · OpenAI SDK
from openai import OpenAI

client = OpenAI(
    api_key="zpr_your_key",
    base_url="https://zapros.ai/api/v1",
)

response = client.chat.completions.create(
    model="openai/gpt-5-nano",
    messages=[{"role": "user", "content": "Hello!"}],
    max_tokens=512,
)
print(response.choices[0].message.content)
JavaScript · fetch
const response = await fetch(
  "https://zapros.ai/api/v1/chat/completions",
  {
    method: "POST",
    headers: {
      Authorization: "Bearer zpr_your_key",
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      model: "openai/gpt-5-nano",
      messages: [{ role: "user", content: "Hello!" }],
      max_tokens: 512,
    }),
  },
);

const data = await response.json();
console.log(data.choices[0].message.content);

Machine-readable references