# AIVAX Documentation > Build, operate, and evaluate AI applications with AIVAX. Language: English. Embedded API references are linked, not fetched. # AIVAX Documentation AIVAX is an AI orchestration platform for building, operating, and evaluating AI applications through one account and API surface. Use hosted or bring-your-own-key (BYOK) models, then add reusable instructions, knowledge, tools, media, user channels, and background processing as your product grows. ## Choose where to start - **Make your first model call:** follow [Getting Started](https://docs.aivax.net/docs/getting-started.md) for a minimal OpenAI-compatible chat completion. - **Understand the platform:** read the [Overview](https://docs.aivax.net/docs/overview.md) to choose between direct inference, AI Gateways, RAG, generations, Batch, and other products. - **Prepare a production integration:** review [Authentication](https://docs.aivax.net/docs/authentication.md), [Pricing](https://docs.aivax.net/docs/pricing.md), and [Plans and limits](https://docs.aivax.net/docs/limits.md). ## Build an AI application - [Inference](https://docs.aivax.net/docs/inference/inference.md) — generate responses with hosted or BYOK models through an OpenAI-compatible API. - [AI Gateways](https://docs.aivax.net/docs/inference/ai-gateway.md) — reuse a model, instructions, RAG, skills, tools, moderation, and inference settings as one assistant runtime. - [RAG collections](https://docs.aivax.net/docs/rag/collections.md) — index your own knowledge for semantic search and grounded answers. - [Rerankers](https://docs.aivax.net/docs/rag/reranking.md) — reorder candidate documents by relevance, with or without a managed collection. - [Skills](https://docs.aivax.net/docs/features/skills.md) — package reusable instructions and operating knowledge for AI Gateways. - [Tools](https://docs.aivax.net/docs/tools/builtin-tools.md) and [MCP](https://docs.aivax.net/docs/tools/mcp.md) — connect assistants to AIVAX capabilities and external systems. - [Chat clients](https://docs.aivax.net/docs/features/chat-clients.md) — publish a gateway through web chat or supported messaging integrations. ## Process text and media - [Text classification](https://docs.aivax.net/docs/rag/classification.md) and [text segmentation](https://docs.aivax.net/docs/rag/text-segmentation.md) — prepare documents for routing, analysis, and retrieval. - [Image generation](https://docs.aivax.net/docs/generations/images.md) — create or edit images from text and reference images. - [Speech generation](https://docs.aivax.net/docs/generations/speech.md) and [audio transcription](https://docs.aivax.net/docs/generations/audio-transcriptions.md) — convert between text and audio. - [Media descriptions](https://docs.aivax.net/docs/generations/media-descriptions.md) — turn images, audio, video, or files into text for downstream processing. - [Voice Sessions](https://docs.aivax.net/docs/inference/voice-session.md) — build low-latency, two-way voice experiences. ## Operate at scale and improve quality - [Batch](https://docs.aivax.net/docs/features/batch.md) — run the same AI workflow over many independent items in the background. - [Agentic Tests](https://docs.aivax.net/docs/inference/agentic-tests.md) — evaluate complete, goal-oriented conversations and track repeatable gateway regressions. - [Structured responses](https://docs.aivax.net/docs/inference/structured-responses.md) — validate generated JSON against an application contract. - [MCP utilities](https://docs.aivax.net/docs/mcp-utilities/account-management-mcp.md) — expose account, collection, documentation, web, and inference capabilities to compatible agents. For endpoint schemas and generated request details, use the [AIVAX API reference](https://inference.aivax.net/apidocs). --- Source: https://docs.aivax.net/docs/overview.html # Overview AIVAX is an AI orchestration platform for building, operating, and evaluating AI applications through one account, API surface, and billing wallet. It combines hosted and bring-your-own-key (BYOK) models with reusable assistant configuration, knowledge retrieval, tools, text and media processing, user-facing channels, background jobs, and conversational evaluation. You do not need every product for every application. Start with direct inference for one response, then add the products that solve a specific reuse, knowledge, integration, scale, or quality requirement. ## Choose the right starting point | Goal | Start with | Why | | --- | --- | --- | | Generate or analyze text in one request | [Inference](https://docs.aivax.net/docs/inference/inference.md) | Call a hosted or BYOK model through the OpenAI-compatible API without creating reusable assistant configuration. | | Reuse instructions, knowledge, tools, and model settings | [AI Gateway](https://docs.aivax.net/docs/inference/ai-gateway.md) | Give your application one stable assistant runtime that can evolve without rebuilding every request. | | Search your own documents or generate grounded answers | [RAG collections](https://docs.aivax.net/docs/rag/collections.md) | Store and index knowledge for semantic retrieval, citations, and gateway context. | | Reorder candidates your application already retrieved | [Rerankers](https://docs.aivax.net/docs/rag/reranking.md) | Improve relevance without requiring a managed AIVAX collection. | | Publish an assistant to end users | [Chat clients](https://docs.aivax.net/docs/features/chat-clients.md) | Connect a gateway to web chat or supported messaging integrations with session and channel controls. | | Process many independent records | [Batch](https://docs.aivax.net/docs/features/batch.md) | Run one repeatable workflow asynchronously with per-item state, validation, retries, cost, and export. | | Test a complete assistant conversation | [Agentic Tests](https://docs.aivax.net/docs/inference/agentic-tests.md) | Simulate a goal-oriented user and judge the gateway across multiple turns. | | Build a low-latency two-way voice experience | [Voice Sessions](https://docs.aivax.net/docs/inference/voice-session.md) | Stream user and assistant audio in an interactive session instead of combining separate audio jobs. | ## Build the assistant runtime ### Inference and AI Gateways AIVAX exposes OpenAI-compatible model listing and chat completion endpoints. Use a **direct model call** for exploration, one-off generation, or configuration that does not need to be reused. Use an **AI Gateway** when the same model, instructions, RAG collections, skills, tools, moderation, or output behavior should serve multiple calls or users. Most production assistants use a gateway because the application can keep calling one identifier while the assistant configuration changes independently. Gateways can use integrated AIVAX models or external OpenAI-compatible providers. Production API base URL: ```text https://inference.aivax.net ``` OpenAI-compatible SDK base URL: ```text https://inference.aivax.net/v1 ``` Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Inference%20(chat%20completions)) ### Knowledge, retrieval, and reranking A [RAG collection](https://docs.aivax.net/docs/rag/collections.md) is a semantic knowledge library. Add documents, test them with [Semantic Search](https://docs.aivax.net/docs/rag/semantic-search.md), and then attach the collection to an AI Gateway when the assistant should answer from that knowledge. AIVAX can also generate grounded answers directly from collections and expose collection search through [Collections MCP](https://docs.aivax.net/docs/mcp-utilities/collections-mcp.md). Reranking is a separate step: it receives a query and candidate documents, then returns the candidates in a more relevant order. Use a collection for managed storage and retrieval; use the standalone [reranking](https://docs.aivax.net/docs/rag/reranking.md) generation when your application already owns the candidates. ### Skills and tools [Skills](https://docs.aivax.net/docs/features/skills.md) package reusable instructions and operating knowledge. Use a skill when the assistant needs to know **how** to perform a task. Use RAG when it needs to retrieve **facts or source material** that may grow or change independently. Tools let the assistant take action or retrieve live information. Choose among: - [Built-in tools](https://docs.aivax.net/docs/tools/builtin-tools.md) for capabilities provided by AIVAX. - [MCP](https://docs.aivax.net/docs/tools/mcp.md) for Model Context Protocol servers and reusable tool ecosystems. - [Protocol functions](https://docs.aivax.net/docs/tools/protocol-functions.md) for HTTP functions defined by your application. - [Shell](https://docs.aivax.net/docs/tools/shell.md) for controlled command execution when the use case requires it. Keep the tool surface as small as the assistant's job allows. Each additional tool expands cost, latency, permissions, and failure paths. ## Process text, documents, and media AIVAX includes focused generation products for work that does not need a full chat conversation: - [Text classification](https://docs.aivax.net/docs/rag/classification.md) assigns labels to one or more documents. - [Text segmentation](https://docs.aivax.net/docs/rag/text-segmentation.md) splits long content into useful chunks for indexing or downstream processing. - [Media descriptions](https://docs.aivax.net/docs/generations/media-descriptions.md) convert images, audio, video, and files into text that another model or workflow can use. - [Image generation](https://docs.aivax.net/docs/generations/images.md) creates or edits images. - [Speech generation](https://docs.aivax.net/docs/generations/speech.md) turns text into audio. - [Audio transcription](https://docs.aivax.net/docs/generations/audio-transcriptions.md) turns audio into text. Use direct multimodal inference when the selected chat model supports the input and should reason over it in the same request. Use a focused generation endpoint when you need a reusable artifact, a transcript, a description, or a preprocessing stage. For large independent input sets, run the appropriate operation through [Batch](https://docs.aivax.net/docs/features/batch.md). For interactive two-way audio, use [Voice Sessions](https://docs.aivax.net/docs/inference/voice-session.md) instead of manually chaining transcription, text inference, and speech generation. ## Deliver, scale, and evaluate ### Chat clients A [chat client](https://docs.aivax.net/docs/features/chat-clients.md) connects an AI Gateway to an end-user channel. It owns presentation, session behavior, allowed origins, uploads, audio replies, channel integrations, and user-facing limits. The gateway continues to own assistant behavior such as the model, instructions, RAG, and tools. Use a chat client for a browser widget or supported messaging integration. Use the inference API directly when your own backend or interface already manages users, conversation state, and delivery. ### Batch [Batch](https://docs.aivax.net/docs/features/batch.md) applies one workflow to dozens or thousands of independent records. A workflow defines the instruction, model or gateway, structured output, validation, tools, and retry policy. A job imports items, processes them in the background, exposes per-item progress and cost, and exports results. Do not use Batch when one item depends on another or when a user needs an immediate answer. Use direct inference for one synchronous result and RAG for searchable knowledge. ### Agentic Tests [Agentic Tests](https://docs.aivax.net/docs/inference/agentic-tests.md) evaluates the configured behavior of an AI Gateway across a bounded conversation. A simulated user pursues a goal while an independent judge evaluates progress. Use persisted tests for reusable, scheduled regression coverage or an ephemeral evaluation for one immediate run. A completed test run is not automatically a successful behavior result. Review the run outcome, judge result, retained conversation, usage, and cost together. ## Operate and connect AIVAX AIVAX records conversations and usage so you can trace behavior, attribute cost, and diagnose failures. The dashboard and account APIs expose account balance, usage, conversations, gateway resources, collection transactions, Batch items, and Agentic Test runs. Start with [Pricing](https://docs.aivax.net/docs/pricing.md) and [Plans and limits](https://docs.aivax.net/docs/limits.md) before enabling a high-volume or media-heavy workflow. AIVAX also provides MCP utilities for compatible agents: - [Account management MCP](https://docs.aivax.net/docs/mcp-utilities/account-management-mcp.md) - [Collections MCP](https://docs.aivax.net/docs/mcp-utilities/collections-mcp.md) - [Documentation MCP](https://docs.aivax.net/docs/mcp-utilities/documentation-mcp.md) - [Web utilities MCP](https://docs.aivax.net/docs/mcp-utilities/web-utilities-mcp.md) - [Media generation MCP](https://docs.aivax.net/docs/mcp-utilities/media-generation-mcp.md) - [Inference MCP](https://docs.aivax.net/docs/mcp-utilities/inference-mcp.md) These utilities expose existing AIVAX capabilities through MCP; they do not replace the underlying account, collection, or inference products. ## Next steps 1. Follow [Getting Started](https://docs.aivax.net/docs/getting-started.md) to make and verify your first chat completion. 2. Read [Authentication](https://docs.aivax.net/docs/authentication.md) before choosing private keys, public keys, or chat sessions for an application boundary. 3. Review [Pricing](https://docs.aivax.net/docs/pricing.md) and [Plans and limits](https://docs.aivax.net/docs/limits.md) before increasing traffic or processing large collections, media, tests, or Batch jobs. 4. Move reusable assistant behavior into an [AI Gateway](https://docs.aivax.net/docs/inference/ai-gateway.md), then add RAG, skills, tools, and a chat client only when the use case requires them. --- Source: https://docs.aivax.net/docs/getting-started.html # Getting Started This guide takes you from an AIVAX account to a verified OpenAI-compatible chat completion. The example uses Python and a private API key from a server-side environment. By the end, you will have confirmed that your key and selected model or AI Gateway can complete a request. ## Before you begin You need: - An AIVAX account with dashboard access and permission to create a private API key. - Python 3.8 or later with `pip` available. For pricing and operational limits, see [Pricing](https://docs.aivax.net/docs/pricing.md) and [Plans and limits](https://docs.aivax.net/docs/limits.md). Production API base URL: ```text https://inference.aivax.net ``` OpenAI-compatible SDK base URL: ```text https://inference.aivax.net/v1 ``` ## 1. Create a private API key Create a **private** key from the API Keys area of the AIVAX dashboard. Copy the key when it is shown and store it as a secret; do not paste the real value into the code below. Private keys are intended for trusted server-side applications. Public keys are restricted credentials for intentionally exposed client-side routes and are not a substitute for a backend key. If you are building a public web widget or messaging experience, review [Chat clients](https://docs.aivax.net/docs/features/chat-clients.md) before exposing any credential. Chat sessions provide a clearer boundary for user identity, conversation history, and attachments. See [Authentication](https://docs.aivax.net/docs/authentication.md) for supported authentication schemes, private and public key behavior, and secret-handling guidance. ## 2. Install the OpenAI SDK Install the SDK in the Python environment you will use for this example: ```bash python -m pip install openai ``` Keep the key outside your source file. For example, set an environment variable named `AIVAX_API_KEY` using the secret-management method appropriate for your shell or deployment platform. ## 3. Choose a model or AI Gateway The `model` field can identify: - A hosted model returned by the model listing endpoint. - An AI Gateway available to your account. Use a **hosted model** for a direct, one-off call or early experiment. Use an **AI Gateway** when you want to reuse the same model, instructions, RAG collections, skills, tools, moderation, and output settings across requests or users. Gateway slugs are supported with private keys. Public-key chat completions must use the full gateway UUID and cannot call integrated models directly. Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Model%20listing) Copy one model name or gateway identifier that is available to your account. You will use it as `` in the next step. ## 4. Make the first request Create a file named `quickstart.py` with the following code: ```python import os from openai import OpenAI client = OpenAI( base_url="https://inference.aivax.net/v1", api_key=os.environ["AIVAX_API_KEY"], ) response = client.chat.completions.create( model="", messages=[ {"role": "user", "content": "Write a one-sentence welcome message."} ], ) print(response.choices[0].message.content) ``` Replace `` with the exact hosted model name or gateway identifier selected in the previous step. Do not replace `AIVAX_API_KEY` with the key itself; the code reads the secret from the environment. Run the file: ```bash python quickstart.py ``` A successful request prints one generated sentence and exits without an API error. Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Inference%20(chat%20completions)) ## 5. Confirm the integration Confirm that the generated response matches the prompt and comes from the model or AI Gateway selected in the previous step. This verifies the endpoint, credential, and model selection used by your application. Before increasing traffic or processing large inputs, review [Pricing](https://docs.aivax.net/docs/pricing.md) and [Plans and limits](https://docs.aivax.net/docs/limits.md). ## Troubleshoot the first request AIVAX uses two response styles: - OpenAI-compatible endpoints return an OpenAI-style `error` object. - Account and administrative endpoints return an AIVAX response envelope with an error or a successful `data` value. | Status | What to check | | --- | --- | | `400 Bad Request` | Confirm the model or gateway identifier and remove unsupported parameters from the request. | | `401 Unauthorized` | Confirm that the private key is present, complete, active, and sent through the SDK configuration. | | `402 Payment Required` | Review [Pricing](https://docs.aivax.net/docs/pricing.md) and confirm that the account is ready for a billable request. | | `403 Forbidden` | Confirm that the key type, model, or selected resource allows this operation. | | `429 Too Many Requests` | Retry later and review [Plans and limits](https://docs.aivax.net/docs/limits.md) before increasing request volume. | | `500 Internal Server Error` | An unexpected AIVAX failure occurred. Retry later; the response does not include internal details. | | `503 Service Unavailable` | A service AIVAX depends on is temporarily unavailable. Retry after the interval in the `Retry-After` header. | If the request still fails, verify in this order: 1. `base_url` is `https://inference.aivax.net/v1`. 2. `AIVAX_API_KEY` is available to the Python process and contains a private key. 3. The selected model or gateway exists and is available to the account. 4. For a gateway, test a plain prompt before adding RAG, tools, media, or structured output so you can isolate configuration problems. ## Long-running inference If a request ends with HTTP `524` or a proxy timeout while AIVAX is still processing it, use the direct inference host: ```text https://direct.inference.aivax.net/v1 ``` The request remains synchronous, not a background job: keep the client connection open until the completion finishes, and configure a client timeout that covers the expected processing time. Use the same private API key, model or AI Gateway identifier, messages, and request parameters. Change the SDK base URL and timeout: ```python import os from openai import OpenAI client = OpenAI( base_url="https://direct.inference.aivax.net/v1", api_key=os.environ["AIVAX_API_KEY"], timeout=300.0, ) response = client.chat.completions.create( model="", messages=[ { "role": "user", "content": "Analyze this case carefully and provide a detailed recommendation.", } ], ) print(response.choices[0].message.content) ``` ## Choose the next product Once the minimal request works, add one capability at a time: - [AI Gateways](https://docs.aivax.net/docs/inference/ai-gateway.md) — make the assistant configuration reusable across requests and users. - [Structured responses](https://docs.aivax.net/docs/inference/structured-responses.md) — require generated JSON to follow an application schema. - [RAG collections](https://docs.aivax.net/docs/rag/collections.md) — index your documents, test retrieval, and attach grounded knowledge to a gateway. - [Built-in tools](https://docs.aivax.net/docs/tools/builtin-tools.md), [MCP](https://docs.aivax.net/docs/tools/mcp.md), or [Protocol functions](https://docs.aivax.net/docs/tools/protocol-functions.md) — let the assistant retrieve live information or take action. - [Chat clients](https://docs.aivax.net/docs/features/chat-clients.md) — deliver a gateway through web chat or supported messaging channels. - [Text and media products](https://docs.aivax.net/docs/overview.md#process-text-documents-and-media) — classify or segment documents, generate images or speech, transcribe audio, and describe media. - [Batch](https://docs.aivax.net/docs/features/batch.md) — apply the same workflow to many independent records asynchronously. - [Agentic Tests](https://docs.aivax.net/docs/inference/agentic-tests.md) — evaluate a complete gateway conversation before and after configuration changes. Before increasing traffic or processing large inputs, review [Pricing](https://docs.aivax.net/docs/pricing.md) and [Plans and limits](https://docs.aivax.net/docs/limits.md). --- Source: https://docs.aivax.net/docs/authentication.html # Authentication AIVAX authenticates API requests with account API keys. AIVAX accepts API keys via: - `Authorization: Bearer ` - `Authorization: Basic ` - `?api-key=` Prefer the `Authorization` header for server-to-server calls. Use the query parameter only when a client or integration cannot send headers, because URLs can be logged by proxies, browsers, and monitoring tools. ## API key types AIVAX has two key families because browser-facing and server-side use cases have different risk profiles. If your code runs on your server, use a private key and keep it out of client bundles, logs, and public repositories. If your code runs in a browser or another environment where the key can be inspected by the end user, use a public key and keep the workflow limited to routes designed for public access. | Type | Prefix | Intended use | Access | | --- | --- | --- | --- | | Private key | `sk-aiv-acc` | Server-side integrations and administrative API calls. | Authenticated account APIs and OpenAI-compatible inference. | | Public key | `pk-aiv-` | Restricted client-side calls to explicitly public routes. | Public RAG query/answer routes and restricted chat-completion calls. | Public keys can be used for RAG semantic search, RAG answer generation, speech generation, media descriptions, image generation, and chat completions. When a public key calls chat completions: - The `model` must be a full AI Gateway UUID; direct integrated-model calls and gateway slug lookup are disabled. - Gateway slug lookup is disabled. - MCP sources, protocol functions, built-in tools, Bash, skills, and sentinel options are stripped from the request. - Only these request parameters are accepted: `model`, `messages`, `prompt`, `temperature`, `top_p`, `top_k`, `seed`, `tools`, `reasoning_effort`, `max_completion_tokens`, `idempotency_key`, and `stream`. - Request and token rate limits are applied both globally per key and per remote address. Use private keys for backend services, account management, model listing, collection management, batch operations, and any workflow that needs the full gateway tool surface. For a first server-side request, continue with [Getting Started](https://docs.aivax.net/docs/getting-started.md). If you are exposing a browser or widget experience to end users, review [Chat Clients](https://docs.aivax.net/docs/features/chat-clients.md) before deciding whether a public key is the right boundary. ## Create and list keys API keys belong to an account and can have a label, expiration, and type. A key with a negative duration does not expire; expired keys are rejected by authentication and are later removed by cleanup jobs. Create separate keys for separate applications. That makes rotation safer: if one integration is compromised, you can revoke only that key instead of breaking every service tied to the account. Labels and expiration dates are operational tools, not decoration; use them to identify who owns the key and when it should be reviewed. Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Create%20API%20Key) [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=List%20API%20Keys) ## Send a key with Bearer auth For OpenAI-compatible SDKs, pass the AIVAX key as the SDK API key and set the base URL to `https://inference.aivax.net/v1`. ```python from openai import OpenAI client = OpenAI( base_url="https://inference.aivax.net/v1", api_key="" ) ``` ## Hook authentication AIVAX can authenticate outbound requests to your services, such as AI Gateway workers and server-side protocol functions. This is the reverse direction of normal API authentication. In a normal API call, your application proves it is allowed to call AIVAX by sending an API key. In a worker or protocol-function callback, AIVAX is calling your service, so your service needs a way to verify that the request really came from the account configuration you control. That is what `X-Request-Nonce` is for. If your account has a hook key, AIVAX sends: ```text X-Request-Nonce: ``` The nonce is a BCrypt hash derived from the account hook key. Validate the header by verifying the stored plain-text hook key against the hash. If the account has no hook key, the header is not sent. Rolling the hook key invalidates existing worker and integration validation secrets. ### C# example ```csharp using BCrypt.Net; var nonce = request.Headers["X-Request-Nonce"]; var hookKey = Environment.GetEnvironmentVariable("AIVAX_HOOK_SECRET"); if (nonce is null || hookKey is null) { return Results.Unauthorized(); } if (!BCrypt.Net.BCrypt.Verify(hookKey, nonce, enhancedEntropy: false)) { return Results.Forbid(); } ``` ### Python example ```python import os import bcrypt from flask import abort, request nonce = request.headers.get("X-Request-Nonce") hook_key = os.getenv("AIVAX_HOOK_SECRET") if nonce is None or hook_key is None: abort(401) if not bcrypt.checkpw(hook_key.encode("utf-8"), nonce.encode("utf-8")): abort(403) ``` ### JavaScript example ```javascript import bcrypt from "bcrypt"; const nonce = req.header("X-Request-Nonce"); const hookKey = process.env.AIVAX_HOOK_SECRET; if (!nonce || !hookKey) { return res.sendStatus(401); } if (!(await bcrypt.compare(hookKey, nonce))) { return res.sendStatus(403); } ``` --- Source: https://docs.aivax.net/docs/pricing.html # Pricing Service usage prices are listed below in USD. **M** means one million tokens; **1k** means one thousand units. Approximate prices (`~`) vary with the model used and the work performed. See [subscription pricing](https://aivax.net/pricing) for monthly plan prices and [Plans and limits](https://docs.aivax.net/docs/limits.md) for quotas. Usage rates are subject to the plan multiplier: - Free: **+25%** on inference taxes; - Pro: **+5%** on inference taxes; - Max: **0%** on inference taxes. BYOK are not affected by inference taxes. Free, Pro, and Max include separate daily allowances for eligible RAG embeddings, Reflex reranking, Julia-1 semantic decisions, and Fetch/OCR extraction. The rates below apply when a metered item is not covered. Coverage is all-or-nothing per item, not necessarily per complete request: an item that cannot fit within the remaining allowance and its permitted margin is billed in full. Compare the allowances and check exclusions in [Plans and limits](https://docs.aivax.net/docs/limits.md#included-daily-subscription-allowances). LLM subscription coverage is currently disabled. ## Inference and Moderation Inference rates depend on the selected model, provider, input size, and media type. Moderation is charged separately in Processing Units (PUs), covering input, cached input, and output usage; its PU price varies with the model and provider used. | Description | Pricing | | --- | ---: | | AI model and AI Gateway inference | Selected model and provider rates | | Input moderation | Variable price per PU; separate from the main inference charge | ## Semantic decisions The decision-model rates below are base USD prices per million input tokens, before account and plan adjustments. Output tokens have no charge in the current decision-model catalog. Julia-1 is eligible for the daily allowance described in [Plans and limits](https://docs.aivax.net/docs/limits.md#included-daily-subscription-allowances); other decision models are billed normally. | Model | Input price / million tokens | | --- | ---: | | `@supersonic-labs/julia-1` | **$0.008** | | `@typesafe/jev-1.13` | **$0.042** | | `@respan/span-01` | **$0.020** | | `@respan/span-01-lite` | **$0.000** | | `@jaredpalmer/kev-4b` | **$0.042** | See [Semantic decisions](https://docs.aivax.net/docs/generations/decisions.md) for model selection and how input usage is measured. ## Agentic Tests Each test includes the selected model or AI Gateway's inference charges, plus simulated-user and judge usage at the selected profile's rates. | Description | Pricing | | --- | ---: | | Model or AI Gateway under test | Regular inference rates | | Low profile - simulated user | Input **$0.25/M tokens**; cache **$0.025/M tokens**; output **$1.50/M tokens** | | Low profile - judge | Input **$0.30/M tokens**; cache **$0.03/M tokens**; output **$2.50/M tokens** | | Medium profile - simulated user | Input **$0.75/M tokens**; cache **$0.075/M tokens**; output **$3.75/M tokens** | | Medium profile - judge | Input **$0.75/M tokens**; cache **$0.075/M tokens**; output **$3.75/M tokens** | | High profile - simulated user | Input **$0.75/M tokens**; cache **$0.075/M tokens**; output **$3.75/M tokens** | | High profile - judge | Input **$1.25/M tokens**; cache **$0.15/M tokens**; output **$4.25/M tokens** | ## RAG and Collections Indexing and search are billed by token usage. Generated RAG responses are charged separately from query embedding, and their price varies with the summarization model. | Description | Pricing | | --- | ---: | | Collection text embedding | **$0.10/M tokens** | | Semantic search - query cache miss | **$0.10/M tokens** | | Semantic search - query cache hit | Zero | | RAG response generation | **~$0.50/M tokens**, excluding query rates | | Reflex - cache miss | **$0.015/M tokens** | | Reflex - cache hit | **$0.003/M tokens** | ## Media Injector Converting media into RAG documents is billed for input, cached input, output, and media usage. The source file, optional context, and generated content affect the total. Rates depend on media type and input-token volume. | Description | Pricing | | --- | ---: | | PDFs and images - up to 272K input tokens | Input **$0.30/M tokens**; cache **$0.03/M tokens**; output **$1.80/M tokens** | | PDFs and images - above 272K input tokens | Input **$0.60/M tokens**; cache **$0.06/M tokens**; output **$3.60/M tokens** | | Audio - up to 256K input tokens | Input/media **$0.60/M tokens**; cache **$0.12/M tokens**; output **$3.00/M tokens** | | Audio - above 256K input tokens | Input/media **$1.20/M tokens**; cache **$0.24/M tokens**; output **$6.00/M tokens** | | Video | Input/media **$0.45/M tokens**; cache **$0.045/M tokens**; output **$3.75/M tokens** | ## Text Tools Text segmentation and classification are billed by token usage. | Description | Pricing | | --- | ---: | | Text segmentation | **$0.30/M tokens** | | Text classification | **$0.10/M tokens** | ## Voice and Media Generation and transcription rates depend on the selected model. Media description pricing is approximate and depends on the available processing model. | Description | Pricing | | --- | ---: | | Voice Sessions | Selected realtime model rates | | Speech-to-text | Varies by model | | Text-to-speech | Varies by model | | Image generation | Fixed output and reference-image tariffs by model | | Media descriptions | **~$1.50/M tokens** | Image generation charges each delivered output at the selected model's fixed output price, plus its per-reference price for every reference sent with that output. Prompt processing is included. Token- and megapixel-priced providers use rounded-up estimates, not exact provider-cost pass-through. No additional AIVAX image-generation markup or account and plan multiplier applies. Current tariffs are listed in the Models catalog; see [Image generation](https://docs.aivax.net/docs/generations/images.md). ## Web Search, OCR and Fetch Web and X searches are billed per search. Advanced web search is billed by token usage and varies with the model and number of interactions. Fetch and OCR extraction use Processing Units (PUs), with a daily free allowance by plan. Optional schema-guided JSON conversion is charged separately. The extraction allowances and PU rates do not apply to JSON conversion or moderation. | Description | Pricing | | --- | ---: | | Web search | **$5/1k searches** | | X (Twitter) search | **$5/1k searches** | | Advanced web search | **~$0.75/M tokens** | | Fetch and OCR extraction - Free | Base daily allowance; uncovered items **$0.15/1k PUs** | | Fetch and OCR extraction - Pro | **10× Free** daily allowance; uncovered items **$0.05/1k PUs** | | Fetch and OCR extraction - Max | **5× Pro** daily allowance; uncovered items **$0.02/1k PUs** | | Fetch JSON conversion (`responseSchema`) | Variable inference-based price per PU; charged separately, with no daily extraction allowance | For the [Fetch API](https://docs.aivax.net/docs/web-foundation/fetch-and-ocr.md), `processingUnits` reports text/OCR extraction usage and `jsonProcessingUnits` reports the additional schema-guided JSON conversion usage. JSON PUs account for input, cached input, and output token usage at the processing model and provider's rates; they are not priced at the plan's OCR rate. The plan's inference multiplier applies to JSON conversion. Omitting `responseSchema` or setting it to `null` disables conversion, reports `jsonProcessingUnits: 0`, and incurs no JSON conversion charge. ## Storage Each plan includes storage. Pro and Max overages are billed hourly at the monthly rates below; Free storage cannot be expanded. | Description | Pricing | | --- | ---: | | Free storage | **30 MB included**; no expansion | | Pro storage | **2 GB included**; excess **$0.50/GB/month** | | Max storage | **20 GB included**; excess **$0.20/GB/month** | ## Other Tools The following tools have no separate tool charge. Model inference used to invoke them is still billed at its regular rate. | Description | Pricing | | --- | ---: | | Memory and calendar | No separate charge | | Advanced requests | No separate charge | | Document generation | No separate charge | | Web page generation | No separate charge | --- Source: https://docs.aivax.net/docs/limits.html # Plans and Limits AIVAX has three account plans: **Free**, **Pro**, and **Max**. The current plan is stored on the account and controls model access, commissions, rate limits, RAG quotas, tool limits, storage quota, conversation retention, and included daily service allowances. For commercial subscription prices and plan packaging, use the [AIVAX pricing page](https://aivax.net/pricing). This page documents the technical limits of the API. ## How limits are enforced Limits are enforced at different layers: - Authentication rejects missing, expired, or unknown API keys. - Public API keys are restricted to public routes and have key-level and per-IP request and token limits. - Balance middleware rejects billable requests when the account balance is below the required minimum. - Storage middleware rejects requests when account storage exceeds the plan quota. - Inference checks model access, request rate, input-token rate, BYOK rate, and Free-plan context size. - RAG checks collection count, search rate, insertion rate, and JSONL import size. - Built-in tools check daily service limits. - Batch processing checks how many workflow items can be processed per day. Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Get%20Account%20Balance) ## Plan limits An em dash (`—`) means the plan does not impose a limit. Model, gateway, provider, or endpoint-specific limits may still apply. | Feature | Free | Pro | Max | | --- | --- | --- | --- | | **Inference** | | | | | Model access | Low-price models | All models | All models | | Inference commission multiplier | 1.25x | 1.05x | 1.00x | | Integrated model requests | 20/min and 500/day | 200/min | — | | Integrated model input tokens | 1,000,000/min | 20,000,000/min | — | | BYOK requests | 30/min | 200/min | — | | Maximum context | 65,536 input tokens | — | — | | LLM subscription coverage | Currently disabled | Currently disabled | Currently disabled | | Standalone text-to-speech requests | 3/min and 40/hour | 30/min | 300/min | | Standalone audio-transcription requests | 3/min and 40/hour | 30/min | 300/min | | Semantic decision requests | 10/min | 50/min | — | | **RAG and collections** | | | | | Collections | 5 | — | — | | Semantic searches | 20/min | 500/min | 3,000/min | | Text-classification documents | 30/min and 300/day | 1,000/min | 10,000/min | | Text-segmentation documents | 10/min and 100/day | 300/min | 2,500/min | | Reranking searches | 30/min | 1,000/min | — | | Reflex processing time | 30 minutes/day | 6 hours/day | — | | Document insertions | 500/day | 10,000/day | — | | JSONL documents per import request | 1,000 | 10,000 | 1,000,000 | | Media Injector | 2 files/day | 30 files/day | 1,000 files/day | | **Built-in tools** | | | | | Web search | 15/day | 1,000/day | 10,000/day | | X/Twitter search | Not available | 1,000/day | 10,000/day | | Advanced web search | Not available | 100/day | 1,000/day | | Document and web page generation | 5/day | 1,000/day | 50,000/day | | Image generation and editing | 5/day | 500/day | 5,000/day | | General service actions | 30/day | 5,000/day | 100,000/day | | Bash commands | 300/hour | 30,000/hour | — | | **Agentic tests** | | | | | New runs per account | 5/min | 30/min | — | | Concurrent runs per account | 1 | 4 | 8 | | **Batch processing** | | | | | Workflow items processed | 500/day | 100,000/day | — | | Files per import request | 1,000 | 1,000 | 1,000 | | Total import size | 100 MiB/request | 100 MiB/request | 100 MiB/request | | Single imported file size | 10 MiB | 10 MiB | 10 MiB | | **Account and support** | | | | | Storage quota | 30 MB | 2 GB | 20 GB | | Cost per excess GB | — | $0.50/GB/month | $0.20/GB/month | | Conversation retention | 2 hours | 2 days | 30 days | | Support level | Email | Priority | Dedicated | ### Semantic decision and Agentic Test rate limits These per-minute limits are shared across API keys belonging to the same account. They are independent of subscription allowances and billing: included usage still consumes the applicable request or run quota. - **Semantic decisions:** each request consumes one unit, regardless of how many questions it contains or which decision model it selects. A request exceeding the account's limit returns `429 Too Many Requests` before evaluation. See [Semantic decisions](https://docs.aivax.net/docs/generations/decisions.md). - **Agentic Tests:** manual runs, scheduled runs, and direct evaluations share one new-run quota. A persisted run consumes its unit when it is queued, not again when execution starts; individual conversation turns do not consume additional run units. Excess manual run requests and direct evaluations return `429 Too Many Requests`. A scheduled test without available quota waits for a later scheduling check rather than creating an extra run. Existing runs remain subject to their separate concurrency and inference limits. See [Agentic Tests](https://docs.aivax.net/docs/inference/agentic-tests.md). Pace requests across the account and use bounded retries with backoff after a 429. An immediate retry still encounters the active rate-limit window. Max has no plan-imposed limit for these two quotas, but other applicable limits remain in effect. ### Semantic decision model limits These model-specific limits apply in addition to the account-level request quotas above. An unspecified context limit does not imply unlimited input. | Model | Context | | --- | --- | | `@supersonic-labs/julia-1` | 1,024 tokens per question | | `@typesafe/jev-1.13` | 32,768 tokens | | `@respan/span-01` | Not specified in the current catalog | | `@respan/span-01-lite` | Not specified in the current catalog | | `@jaredpalmer/kev-4b` | 8,192 tokens | Julia-1 has additional serving limits: | Limit | Value | | --- | --- | | Questions per request | 1–32 | | Choices or score levels per question | 2–20 | | Boolean options | Exactly two: false and true | | Combined context per question | 1,024 tokens, including state, question, options, and special tokens | | Question and options budget | 256 tokens within the combined context | | Description of an individual option | At most 48 tokens | | Decision payload limit | 256 KiB | These limits interact: twenty options can exceed the combined question/options budget even if each description fits its individual limit. The current AIVAX Julia-1 serving limits apply even if an upstream model card lists a larger context. See [Semantic decisions](https://docs.aivax.net/docs/generations/decisions.md) for usage and error guidance. ### Included daily subscription allowances Free, Pro, and Max include separate daily allowances for the services below. Each comparison refers to the same service on the named plan, not to a shared credit balance or a guaranteed number of requests. Unused allowance from one service cannot cover another. Reseller accounts do not receive subscription allowances. | Included service | Free | Pro | Max | | --- | --- | --- | --- | | RAG search and insertion embeddings | Base allowance | 25× Free | 4× Pro | | Reranking with Reflex | Base allowance | 5× Free | 10× Pro | | Semantic decisions with Julia-1 | Base allowance | 2.5× Free | 2× Pro | | Fetch and OCR extraction | Base allowance | 10× Free | 5× Pro | RAG searches and document insertions share the embedding allowance. It does not cover answer generation, media processing, text classification, or segmentation. A query embedding served from cache does not consume it. Reflex uses a separate reranking allowance that includes both cached and uncached input. Julia-1 is currently the only decision model covered by the semantic decision allowance; other decision models are billed normally. Optional Fetch JSON conversion is separate from the extraction allowance. Coverage is evaluated for each metered service item: a document's embedding, an individual query-term embedding, a reranking call, a decision call's input usage, or an extraction operation. Each item is either fully included or billed in full at normal rates. Included items are tracked in subscription consumption, not as zero-cost entries in billing history. The current allowances permit a 10% margin above their base capacity. An item that would exceed that margin leaves the allowance unchanged and is billed normally. One request can contain several items, so some may be included while others are charged. Daily allowances reset at midnight in the server's local time. Check the account's subscription usage indicators for consumption and reset status; usage can exceed 100% while within the margin. LLM subscription coverage is currently disabled, so text-model inference and RAG answer generation remain metered separately. Allowances do not bypass balance requirements, rate limits, or Reflex's separate processing-time cap. See [Pricing](https://docs.aivax.net/docs/pricing.md) for charges when an item is not covered. Reseller accounts support 8 concurrent agentic test runs per account. Integrated model requests are limited by both request count and input tokens. Model rate-limit groups adjust the request-count thresholds: | Rate-limit group | Threshold multiplier | | --- | --- | | Common | 1.0x | | Discounted | 0.5x | | Low | 0.3x | | Free | 0.1x | For example, a Pro account normally has 200 integrated-model requests per minute. With a `Discounted` model group, the adjusted threshold is 100 requests per minute. BYOK uses a provider key configured on the gateway instead of an integrated AIVAX model, but requests still pass through AIVAX infrastructure and use the plan's BYOK limit. Text-classification and text-segmentation quotas each count every item in the request's `documents` array, not each HTTP request. A request that would exceed any active window returns `429 Too Many Requests`. Text classification uses the default embedding model and is billed for the embedding work performed. Standalone text-to-speech and audio transcription each use their own plan request quota. Voice Sessions use the selected realtime model and are subject to applicable model access, balance, and inference limits instead of these standalone request quotas. Input transcription is not currently supported inside Voice Sessions. The JSONL import endpoint rejects a request when it reaches the plan's per-request document limit. The reranking limit applies to the autonomous reranking endpoint and to RAG searches that use a reranker, including searches performed through AI Gateways and MCP tools. The Reflex limit counts the time spent processing Reflex requests. It applies to the autonomous reranking endpoint and RAG searches that use Reflex; cached input does not consume the quota separately. Requests that exceed the plan limit return `429 Too Many Requests`. See [Reflex](https://docs.aivax.net/docs/rag/reflex.md) for request limits, cache behavior, and pricing. General service actions share the service-action quota shown above. Batch processing is asynchronous; if processing is paused or fails because of quota, retry after the quota window resets or upgrade the account. ## Public API keys Public keys have additional limits independent of the account plan. | Scope | Request limits | | --- | --- | | Per remote address | 3/5s, 20/min, 300/hour, 1,000/day | | Global per key | 10/5s, 60/min, 1,500/hour, 10,000/day | | Scope | Token limits | | --- | --- | | Per remote address | 100,000/5min, 500,000/30min, 2,000,000/6h, 5,000,000/day | | Global per key | 500,000/5min, 2,000,000/30min, 10,000,000/6h, 25,000,000/day | Public keys can be used for RAG semantic search, RAG answer generation, speech generation, media descriptions, image generation, and chat completions. For chat completions, public keys also require a full AI Gateway UUID, restrict request parameters, and omit server-side tool surfaces. See [Authentication](https://docs.aivax.net/docs/authentication.md). --- Source: https://docs.aivax.net/docs/data-collecting.html # Data Collecting AIVAX offers an optional semantic data collection program for accounts that choose to contribute eligible RAG and reranking data to model development. Document indexing and storage are outside this program. All collected RAG and reranking records are anonymized before they are written to the training dataset. The setting is disabled by default and must be enabled by an authorized Account Manager — the person using the AIVAX account, not a role in the API. ## What changes when collection is enabled While the setting is enabled: - eligible RAG query-embedding usage receives a 10% discount; - eligible reranking operations receive a 10% discount; and - document indexing, storage, unrelated inference, tools, and other services keep their regular prices. The discount applies only to eligible operations performed while collection is enabled. Disabling collection removes the discount from future operations. ## Data included Depending on the operation, AIVAX may collect: | Operation | Collected data | | --- | --- | | RAG semantic search | Query terms, selected reranker, and the contents and relevance scores of returned documents. | | Reranking | Query, submitted documents, selected reranker, and ranking results. Unknown response metadata and usage or billing information are not included. | The anonymized dataset does not store the account ID, API key, request ID, collection or document IDs, document names, billing data, or collection timestamp. The account is consulted transiently only to verify that collection is enabled and to apply the eligible operation discount. Document indexing and stored collections are never copied into the training dataset by this program; document content is included only when it is submitted for reranking or returned by a RAG search. Enabling this setting does not by itself include unrelated chat completions, conversations, tool calls, or other account resources in the training dataset. ## Purpose and use AIVAX may use collected data to develop, train, fine-tune, evaluate, test, and improve models and systems related to embeddings, retrieval, ranking, reranking, and other semantic processing. This may include preparing datasets, annotating or transforming records, measuring quality, and producing aggregated or derived artifacts. Access is limited to authorized personnel and service providers that support these purposes under applicable confidentiality and data-protection obligations. AIVAX does not sell collected semantic data or maintain an account-to-record mapping for this dataset. ## Account Manager responsibilities In AIVAX legal terms, the Account Manager is the person using the AIVAX account. Before enabling collection, the Account Manager must: - have authority to accept these conditions for the account; - have an appropriate legal basis for AIVAX to use the submitted data for the purposes above; - provide any notices and obtain any permissions or consents required from end users or other data subjects; and - avoid submitting credentials, secrets, regulated data, or sensitive personal data unless its collection and use are legally permitted and necessary. ## Enabling, disabling, and deletion The Account Manager can control collection in **Dashboard > My account > Semantic data collection**. - The setting is disabled by default. - Enabling it authorizes anonymized collection of future eligible RAG and reranking operations. - Disabling it stops new collection and ends the discount for future operations. - Disabling it does not automatically delete records collected while consent was active or reverse training already completed. Because account identifiers and account-to-record mappings are not stored, AIVAX cannot retrieve or delete training records merely from an account ID. Requests concerning personal data present inside submitted semantic content can be sent to **privacy@aivax.net** or **wm@aivax.net** and must include enough information to locate the content where applicable. Source records are retained only for as long as reasonably necessary for the documented purposes, legal obligations, security, and audit requirements, and may then be deleted. Deletion of source records does not require AIVAX to retrain or destroy models or aggregated artifacts that no longer identify a person, except where required by applicable law. See the [Privacy Policy](https://docs.aivax.net/docs/legal/privacy-policy.md) and [Terms of Use](https://docs.aivax.net/docs/legal/terms-of-service.md) for the governing legal terms. --- Source: https://docs.aivax.net/docs/changelogs.html # Changelogs Technical changes that affect AIVAX products, services, or the public API. Dates identify when entries were added or updated, not confirmed production rollout dates. Each item identifies the affected product or service; maintenance with no user-facing effect is omitted. ## Saturday, October 3rd, 2026 Changes: - **Documentation — Updated reading and search experience.** Documentation adds language-specific search, light and dark themes, mobile navigation, and a page menu for viewing or copying Markdown. AI agents can use the documentation index and full-text files. Existing documentation URLs continue to resolve; some older guides redirect to their current product documentation. ## Friday, October 2nd, 2026 Changes: - **Models — Claude Sonnet 5.5 and GPT-6.1 Sol added.** Adds `@anthropic/claude-5.5-sonnet` (Claude Sonnet 5.5), the direct successor to Claude Sonnet 5, and `@openai/gpt-6.1-sol` (GPT-6.1 Sol), an upgrade to GPT-6 Sol, to the text model catalog. Both support thinking, image and file input, and tool calling; GPT-6.1 Sol also supports structured output. The `@model-router/claude:mid` and `@model-router/openai:mid` aliases now select Claude Sonnet 5.5 and GPT-6.1 Sol, replacing Claude Sonnet 5 and GPT-6 Sol, respectively. Applications using these aliases may see changes in response quality, latency, and cost. Existing explicit model identifiers remain unchanged. ## Thursday, October 1st, 2026 Breaking changes: - **API — Error status codes reflect the failure source.** Failures are no longer all returned as `400 Bad Request` by account, RAG, web, and generation endpoints. Invalid requests still return `400`, and authentication, balance, and permission failures that previously surfaced as `400` now return `401`, `402`, or `403`. Unexpected AIVAX failures return `500 Internal Server Error` with a generic message, and unavailable external services, including model providers, return `503 Service Unavailable` with a `Retry-After` header. On OpenAI-compatible endpoints, the error `code` is now `invalid_request_error`, `service_unavailable`, or `server_error` instead of always `server_error`. Clients that treat every non-2xx response as a request error should retry `500` and `503` responses, honoring `Retry-After`. See [Troubleshoot the first request](https://docs.aivax.net/docs/getting-started.md#troubleshoot-the-first-request). Fixes: - **Chat completions — Malformed tool declarations rejected as invalid requests.** Requests whose `tools` entries have missing or wrongly typed fields now return `400 Bad Request` instead of a server error. - **Gateways — Bash tool reports invalid options to the model.** Unknown options or values that do not match a tool's parameters now return a command error that the model can correct, instead of failing the tool call. Changes: - **MCP utilities — Media generation MCP.** A new hosted MCP server at `https://inference.aivax.net/v1/mcp/media-generation` exposes `list_models`, `generate_image`, and `generate_speech` to MCP-compatible clients. `list_models` takes a `type` of `image` or `audio`; generated images and MP3 speech are returned as public URLs. Use the optional `X-Mcp-Enabled-Tools` header to choose which tools a client sees. Generations use the same pricing and limits as the Image Generation and Speech Generation APIs and require a positive balance. See [Media generation MCP](https://docs.aivax.net/docs/mcp-utilities/media-generation-mcp.md). ## Monday, September 28th, 2026 Fixes: - **Gateways — Bash tool help accepts nullable parameters.** Requesting help for tools whose parameters accept multiple types, including `null`, no longer fails while listing their arguments. Help preserves the accepted types and nested parameter paths. No tool schema changes are required. - **Chat integrations — Failure notifications restored.** Streaming and non-streaming conversations again attempt to send “System: something went wrong. Please, try again later.” after an unrecoverable generation failure, exhausted recovery attempts, or a turn that sends no message. A final failure is reported even if an earlier partial reply was delivered. Notification delivery still depends on the messaging service being available. Changes: - **Gateways / MCP — Server instructions and remote skills.** MCP sources can include server instructions and root skill documents alongside tools, including sources added by workers. Both options default to enabled; set `allowClientInstructions` or `allowRemoteSkills` to `false` to exclude that content. Skill discovery requires the remote server to advertise compatible capabilities; account skills remain available. Remote skill content is checked against its advertised size, digest, and frontmatter. Dynamic skills, supporting files, scripts, and servers requiring the newer discovery protocol are not supported. See [MCP functions](https://docs.aivax.net/docs/tools/mcp.md#server-instructions-and-remote-skills) for compatibility and trust limits. - **Gateways — Time zone selection.** The current date and time setting now offers a time zone dropdown grouped by region, including UTC. Existing saved time zones are preserved when reopening the setting. - **Models — Seven new text models.** Adds `@cohere/command-a-plus`, `@upstage/solar-mini4`, `@aion-labs/aion-3.5`, `@aion-labs/aion-3.5-mini`, `@qwen/qwen3.8-max-prime`, `@z-ai/glm-5.3-prime`, and `@fireworks/ember-1` as selectable text models, billed per provider at published token rates with existing account and plan adjustments still applying. Upstage and Fireworks models now display their provider icons instead of the generic fallback. Existing model identifiers remain unchanged. ## Sunday, September 27th, 2026 Changes: - **Telegram — Compact tool progress.** Streamed replies show only the latest tool preamble in the thinking indicator when tool-call visibility is enabled, instead of accumulating tool-name blocks in the response. The final reply contains no tool-progress blocks. Other messaging channels and non-streamed replies are unchanged. - **Gateways — Bash tool selection.** The Bash tool list now includes a shortcut for `get_date_time` (Current date and time). Inclusion and exclusion lists accept case-insensitive wildcard patterns: `*` matches any number of characters, as in `something_*`, and `?` matches one character. Exact tool names remain supported. - **Semantic decisions and Agentic Tests — Account rate limits.** Per-minute account limits apply to semantic decision requests and new Agentic Test runs; manual, scheduled, and direct evaluations share the test-run limit. Requests above the limit return HTTP 429, while scheduled tests wait for a later scheduling check. Clients should pace requests and retry after the rate-limit window clears. Existing inference limits still apply. See [Plans and Limits](https://docs.aivax.net/docs/limits.md#semantic-decision-and-agentic-test-rate-limits), [Semantic decisions](https://docs.aivax.net/docs/generations/decisions.md#account-rate-limits), and [Agentic Tests](https://docs.aivax.net/docs/inference/agentic-tests.md#run-and-inspect-a-test). - **Models — Consistent speech synthesis prices.** Text-to-speech catalog prices now derive from the same per-character rates used to calculate synthesis usage, displayed per 1,000 characters. Existing model identifiers and billing rates remain unchanged. - **Models — Comparable transcription prices.** Speech-to-text model prices are displayed consistently in USD per minute, converting hourly and per-second rates for comparison without changing billing rates or duration measurement. - **Image generation — Fixed output and reference prices.** Image generation uses a fixed price per delivered output plus a per-reference price for each output. Models whose providers bill tokens or megapixels now use rounded-up estimates instead of metered token charges. Prompt processing is included in the output estimate; models without a separate reference charge list a zero reference rate. These are fixed tariffs, not receipts for the provider's actual consumption. Image charges no longer include the AIVAX image-generation markup or account and plan pricing multipliers. See [Image generation](https://docs.aivax.net/docs/generations/images.md). - **Models — Subscription coverage.** The Models page now includes a “Subscription Usage” column for rerankers and semantic decision models. “Included” identifies models eligible for daily plan allowances; “-” identifies models without coverage. Eligibility does not indicate an account's remaining allowance. Both service catalogs expose this eligibility as `subscriptionUsage`. - **Subscriptions — Included daily usage allowances.** Free, Pro, and Max subscriptions include separate daily allowances for RAG search and insertion embeddings, Reflex input (including cached input), and Julia-1 semantic decisions. Each metered item is either fully included or billed at normal rates; a request with multiple items can combine included and paid usage. This also applies to OCR extraction, replacing partial coverage. Included items appear in subscription consumption without zero-cost billing-history entries; usage indicators may exceed the base allowance within the permitted margin. Uncovered items retain normal billing records. `includesSubscriptionModels` is false while inference subscriptions are disabled; LLM subscription coverage remains disabled. Reflex's daily processing-time cap remains separate, and reseller accounts do not receive subscription allowances. See [Plans and Limits](https://docs.aivax.net/docs/limits.md#included-daily-subscription-allowances) and [Pricing](https://docs.aivax.net/docs/pricing.md). ## Saturday, September 26th, 2026 Breaking changes: - **Image generation — Deprecated models removed.** Removes `majicMIX-realistic`, `AbsoluteReality`, `CyberRealistic`, `CyberRealistic-Pony`, `RealCartoon-Realistic`, `Hassaku-XL`, and `Meina-Mix` from the available models. Direct API requests using these identifiers now fail; select an active model instead. Built-in image generation uses `flux-schnell` when no valid model is configured, replacing the deprecated `AbsoluteReality` default. Review saved configurations; output style and pricing differ. Changes: - **Image generation — Additional Pollinations models.** Adds 16 official raster-image models, including FLUX 1.1 Pro and FLUX 2 variants, MAI Image variants, GPT Image 2.5 Flare and Sunburst, Qwen Image 2.1 and 3, Grok Imagine Image 2.0, Recraft V4.1 Flash, Krea 2 Medium, DreamShaper 8 LCM, and Seedream 5 Pro, with model-generated previews in the image picker. Community and SVG models are excluded. The catalog displays the billing units. See the September 27th entry for the subsequent fixed-price billing change. Failed generations are not counted as delivered images. See [Image generation](https://docs.aivax.net/docs/generations/images.md). - **Models — Service model catalogs.** The Models page now includes tables for image generation, speech-to-text, text-to-speech, reranking, and semantic decisions, with backend-provided descriptions, release dates, base USD pricing with billing units, and an Actions dropdown in every service-model table for copying model names and opening integration documentation. Model names show a friendly label when available while copying the identifier accepted by the API. Pricing uses compact input, cached-input, output, image, character, and duration units separated by arrows where applicable, with full rates and units available in the tooltip. Catalogs are ordered newest first. OpenRouter catalog dates and friendly names supplement missing launch metadata; catalog dates are explicitly labeled rather than presented as manufacturer release dates. Models without either date remain last. Failed catalog requests can be retried independently. The public information catalogs expose these details, including the new `GET /api/v1/information/speech-models.json` catalog. Account and plan pricing adjustments still apply. See [Pricing](https://docs.aivax.net/docs/pricing.md) and [Semantic decisions](https://docs.aivax.net/docs/generations/decisions.md). - **Privacy — Judicial disclosure and retention clarified.** The [Privacy Policy](https://docs.aivax.net/docs/legal/privacy-policy.md) and [Terms of Use](https://docs.aivax.net/docs/legal/terms-of-service.md) specify Brazilian court orders for disclosure, foreign requests to preserve existing logs for up to 1 year, and up to 1 year of technical logs and metadata. They clarify that conversation content is collected only when Conversations is enabled for the request or account, and that available account resources and backups dating back up to 3 months may be disclosed under a Brazilian court order. The content license in the Terms is expressly subject to these limits. - **Generations — Additional decision models.** The [Semantic decisions guide](https://docs.aivax.net/docs/generations/decisions.md) explains question types and response interpretation. The public `GET /api/v1/information/decisions-models.json` endpoint lists canonical names, aliases, supported question types, context lengths, release dates, and base token prices. Adds `@respan/span-01`, `@respan/span-01-lite`, `@jaredpalmer/kev-4b`, and `@supersonic-labs/julia-1` as model choices for semantic decisions. Julia-1 supports `choice`, `score`, and `noul`; its input usage includes the state repeated for each question. Existing account and plan pricing adjustments apply, and model identifiers and request formats remain unchanged. See [Plans and Limits](https://docs.aivax.net/docs/limits.md#semantic-decision-model-limits) and [Pricing](https://docs.aivax.net/docs/pricing.md#semantic-decisions) for current limits and rates. ## Tuesday, September 22nd, 2026 Changes: - **Models — GPT-6 Sol and Luna added.** Adds `@openai/gpt-6-sol` and `@openai/gpt-6-luna`, including their `:pro` reasoning variants. The `@model-router/openai:mid` and `@model-router/openai:budget` aliases now select GPT-6 Sol and GPT-6 Luna, respectively. Applications using these aliases may see changes in response quality, latency, and cost. Existing explicit model identifiers remain unchanged. ## Monday, September 21st, 2026 Breaking changes: - **Gateways — Off-topic moderation is being removed.** The dedicated off-topic threshold will no longer block requests that stray from the conversation's purpose. If your application relies on this check, review its topic restrictions before adopting this change; the remaining moderation categories are not an equivalent replacement. - **Built-in tools — Advanced web research is being disabled.** The `AdvancedWebUsage` tool will return an unavailable response instead of performing research. Remove reliance on this tool from gateway instructions and workflows. Standard [Web Search](https://docs.aivax.net/docs/web-foundation/web-search.md) and URL extraction remain separate alternatives, not equivalent replacements for multi-step research. - **Models — Mercury 2.5 model identifier changed.** Replace `@inception/mercury-2.5-preview` with `@inception/mercury-2.5` in requests and gateway settings. The preview identifier is no longer listed, and the replacement is no longer marked as preview. - **Gateways / MCP — MCP tool names are source-qualified.** Tools from different MCP sources no longer share an unqualified name in the gateway. Review gateway instructions, tool-selection rules, and workers that match exact tool names. The original tool name at the connected MCP server is unchanged. See [MCP functions](https://docs.aivax.net/docs/tools/mcp.md). Fixes: - **Collections, Gateways, and Generations — Service model availability.** Collection answer generation, gateway routing, chat utilities, and media-processing services avoid selecting temporarily unavailable models. Teach Skill and multimodal preprocessing can try another available model after a retryable failure; success still depends on service availability. - **Teach Skill — Usage calculation.** Teach Skill processing charges use the model pricing associated with the completed request, including when a retry changes the model used. See [Teach Skill](https://docs.aivax.net/docs/generations/teach-skill.md) and [Pricing](https://docs.aivax.net/docs/pricing.md). - **Fetch and OCR — More reliable page extraction.** Cancelled or timed-out page-content requests no longer leave page extraction running indefinitely. This addresses cases where web content extraction could stall. - **Chat clients — Final reply and message history.** `completionText` now selects the final generated assistant reply rather than combining it with earlier assistant text during tool use. The added `createdMessages` response field preserves newly generated messages in order, including tool interactions; submitted messages are not repeated. Use this field when you need the complete generated turn. - **Gateways / MCP — Structured MCP results are retained.** Assistants receive structured result content in addition to supported content blocks, avoiding missing information when an MCP tool returns structured output. - **Fetch and OCR — X post extraction.** Improved readable-text extraction from public X post links. Content availability and access restrictions still apply. - **Fetch and OCR — Plain-text normalization changed.** Plain-text conversion no longer collapses internal whitespace or normalizes Unicode characters. Surrounding whitespace may still be trimmed in extraction results. Applications that compare extracted text exactly or require normalized spacing should perform that normalization themselves. Changes: - **Fetch and OCR — Optional structured extraction.** Supply `responseSchema` to convert extracted content into JSON. Results add `extractedObject` and `jsonProcessingUnits` while retaining `extractedText` on success. Without a schema, the new fields are null and zero respectively. JSON conversion is billed separately from extraction and is not covered by the daily extraction allowance. See [Fetch and OCR](https://docs.aivax.net/docs/web-foundation/fetch-and-ocr.md). - **Generations — Semantic decisions.** Evaluate multiple named questions against a shared JSON state using `noul` (true/false criteria), `choice`, or `score`. Responses include named answers, token usage, and cost. The service requires a positive balance, and its token usage contributes to account usage totals. - **Built-in tools — Current date and time.** The `DateTime` option lets assistants request the current date, time, weekday, time zone, and UTC offset. Set `dateTimeTimeZone` to the desired time zone; the default is `America/Los_Angeles`, with daylight-saving adjustments, independent of the browser's time zone. Invalid time zones are rejected. - **Models — Additional model choices.** Adds GLM-5.3-FlashX, Fugu Max, Pareto, Ling 3.0 Flash VL, MiMo-V2.6-Pro, MiMo-V2.6-Flash, MiMo-V2.6-Pro-UltraSpeed, and Grok 4.7 to the inference catalog. The Xiaomi frontier, mid, and budget aliases now select MiMo V2.6 models, while `@model-router/grok:latest` selects Grok 4.7. Alias users may see different response quality, latency, and cost; availability and supported capabilities depend on the selected model and account plan. - **Inference — Transient-failure recovery.** Inference requests can make additional recovery attempts when a provider is temporarily unavailable. This may avoid some failed requests, but can also increase response time before an error is returned. - **Avi Assistant — Updated default model.** The AIVAX console assistant changes its default model, which may change response style, latency, and usage cost. This does not change the model selected in your own gateways. - **Documentation — Updated service guidance.** The API documentation overview and Avi Assistant guidance cover voice sessions, agentic tests, classification, segmentation, reranking, and web extraction, with current documentation links and service-specific billing guidance. This is a guidance update, not the introduction of those services. - **Models — DeepSeek V4.1 Flash added.** The inference catalog includes `@deepseek/deepseek-v4.1-flash` with tool-calling support. Check model availability and plan eligibility before selecting it. - **Models — Router selections updated.** `@model-router/deepseek:latest` and `@model-router/deepseek:budget` now select DeepSeek V4.1 Flash. Applications using these aliases may see different response quality, latency, and cost without changing the alias. Adds `@model-router/claude:frontier-mythos` and `@model-router/mercury:latest` as additional choices. - **Chat clients — Structured prompt input.** Synchronous chat-client prompts accept plain text, a single message, or an ordered array of messages, including assistant tool calls and matching tool results. Existing single-message input remains supported. Optional `instructions` adds context for that request without replacing saved session context. See [Chat clients](https://docs.aivax.net/docs/features/chat-clients.md). - **Chat clients — Turns without session persistence.** Set `commit` to false to generate a reply without saving the submitted and generated messages to session history. The default remains true. This is not a free preview: inference and tool actions still run. To continue an uncommitted tool interaction, send the assistant tool-call message together with its tool results. - **Agentic Tests — Optional testing notifications.** Account notification preferences can enable weekly test summaries and alerts when a test reaches three consecutive failures. These complement existing failure and recovery notifications. See [Agentic Tests](https://docs.aivax.net/docs/inference/agentic-tests.md). - **Fetch and OCR — Rendered web content.** HTML extraction supports rendered page content, improving coverage of pages whose readable text depends on scripts. Rendering is metered in processing units; this does not guarantee access to every website or restricted page. See [Fetch and OCR](https://docs.aivax.net/docs/web-foundation/fetch-and-ocr.md). --- Source: https://docs.aivax.net/docs/rag/collections.html # Collections and Documents AIVAX provides a RAG (Retrieval-Augmented Generation) service for storing documents and retrieving them later through semantic search. A collection is an account-owned group of documents. Each document stores text, optional tags, an optional reference, optional metadata, and the vectors generated by the indexing job. Collections can be searched directly through the RAG API or attached to an AI Gateway so retrieved documents can be injected into the model context. Compare vector storage with the surrounding ingestion and retrieval pipeline in [RAG vs vector database](https://aivax.net/blog/a-vector-database-is-not-a-rag-system/). ## Collections Use collections to group documents that belong to the same knowledge base, product, tenant, language, or operational purpose. A collection is the container you create before adding searchable knowledge. Think of it as the boundary for a knowledge base: a support collection can hold help-center answers, a legal collection can hold contract clauses, and a product collection can hold descriptions, policies, and troubleshooting notes. Later, you can search the collection directly with the [Semantic Search](https://docs.aivax.net/docs/rag/semantic-search.md) API, expose it through [Collections MCP](https://docs.aivax.net/docs/mcp-utilities/collections-mcp.md), or attach it to an [AI Gateway](https://docs.aivax.net/docs/inference/ai-gateway.md) so retrieved documents are placed into the model context automatically. Each collection has: - A unique collection ID. - A name. - Optional context and contextual tags. - A set of documents. - Usage statistics based on RAG transactions. Collection availability and account limits depend on the current account configuration. See [Plans and limits](https://docs.aivax.net/docs/limits.md) before creating collections for production use. ## Documents A document is the unit that gets indexed and retrieved. It should be small enough to match a specific question and complete enough to be useful on its own. This is the part that most affects RAG quality. A document should not be "everything you know" about a source; it should be one piece of knowledge that can stand alone when the model reads it later. If a user asks about cancellation fees, the retrieved document should already contain the relevant rule, product, condition, and exception. If the answer only makes sense when the model also sees the previous page, the document is probably too dependent on surrounding context. A good document usually has: - A stable name. - Focused text. - Optional tags for filtering or maintenance. - Optional metadata for application-specific data. - An optional reference ID when the document is one chunk of a larger logical item. For example, a car manual should not be indexed as one document. Index separate documents for topics such as starting the vehicle, checking tire pressure, pairing Bluetooth, and replacing a headlight. Each document should include enough context to be read independently. For broader chunking guidance, see [Best Practices for RAG](https://docs.aivax.net/docs/rag/best-practices.md); for query behavior after indexing, see [Semantic Search](https://docs.aivax.net/docs/rag/semantic-search.md). ## Document Fields When you import documents in JSONL, each line represents one document that can be created or updated. The important field is `docid`: it is the stable name AIVAX uses to recognize the same document on future imports. If you send the same `docid` again with different text, the existing document is updated and reindexed. If you only need to preserve extra application data, use `__meta` instead of mixing that data into the searchable text. The JSONL import endpoint accepts one JSON object per line: | Property | Type | Required | Description | | --- | --- | --- | --- | | `docid` | `string` | Yes | Stable document name. Existing documents are matched by this value. | | `text` | `string` | Yes | Text content to index semantically. | | `__ref` | `string` | No | Reference ID used to group related chunks. Maximum stored length is 64 characters. | | `__tags` | `string[]` | No | Tags for filtering, browsing, and maintenance. | | `__meta` | `object` | No | Metadata returned with document details and search results. Metadata is not the semantic text used for embeddings. | The document name must be non-empty and is limited by the API to 256 characters. The stored document content is required and cannot be empty. ## Upserts and Reindexing Documents are matched by name (`docid` in JSONL, `Name` in the single-document API). When a document is created, it is queued for indexing. When existing document text changes, the document is queued again and its vectors are regenerated by the background indexer. When only `__meta` changes, metadata is updated without reindexing the document text. Reference and tag values are stored with the document. In the single-document API, changed reference, tags, or metadata can update an existing document without reindexing when the text is unchanged. In the JSONL import endpoint, changed text queues reindexing, metadata-only changes update metadata without reindexing, and changed text can also update reference, tags, and metadata. A JSONL entry that changes only the reference or tags is skipped. ## References Use `__ref` when multiple documents represent parts of the same logical source, such as: - Sections of the same contract. - Clauses from the same policy. - Chunks from the same PDF. - Product fragments that should be shown together. When search reference expansion is enabled, if one chunk matches, other documents in the same collection with the same reference can be included in the response. ## Media File Import The AIVAX dashboard can upload a source file and process it into RAG documents with [Media Injector](https://docs.aivax.net/docs/rag/media-injector.md). Use it when you have a source file but do not already have focused, self-contained document text prepared for direct or JSONL import. The original file name is normalized to Unicode NFC and preserved during upload, including accented letters, non-Latin scripts, typographic punctuation, and other Unicode characters. You do not need to rename a file to an ASCII-only name before importing it. A Media Injector job is created only after every file chunk has uploaded and the dashboard successfully completes the upload. You can then follow it under **Batch > Media Processing**. If no job appears, the upload did not reach its completion step; retry the upload and check the error shown by the dashboard. ## Batch Import Limits Batch import is sent as a JSONL file in the `documents` multipart field. Use batch import when you already have many documents prepared outside AIVAX, such as chunks generated from PDFs, product catalogs, policies, or help-center articles. If you are creating or updating one document from an application flow, the single-document endpoint below is usually easier. If you are preparing a large knowledge base, import in batches, wait for indexing, and then test retrieval through [Semantic Search](https://docs.aivax.net/docs/rag/semantic-search.md) before attaching the collection to a production gateway. Per-request JSONL line limits and daily RAG insertion limits vary by plan; see [Plans and limits](https://docs.aivax.net/docs/limits.md#plan-limits). If your import exceeds the request limit, split it into multiple files. If your account reaches the daily insertion limit, wait for the rate window to reset or upgrade the plan. > [!WARNING] > Indexing incurs cost based on document text tokens when documents are created or when their text changes. The embedded API reference is the source of truth for the import request and response contract. [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Index%20Documents%20(JSONL)) ## Document Management ### Create or update document This endpoint is useful when your application manages documents one at a time. For example, an admin screen can save one FAQ entry, one policy clause, or one product note directly into a collection. AIVAX matches the document by name: changed text queues reindexing, while metadata-only changes update metadata without reindexing the content. [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Create%20or%20Update%20Document) ### List documents The browse endpoint helps you inspect what is already inside a collection. Use it when you need to verify an import, find a document by name, review queued versus indexed documents, or filter content before deciding whether to update, delete, or reimport part of the knowledge base. Supported filters: - `-t "tag"`: documents containing the tag. - `-r "reference"`: documents with the exact reference ID. - `-c "content"`: documents whose content contains the text snippet. - `-n "name"`: documents whose name contains the text snippet. - `-i "id"`: documents whose ID contains the supplied text. Supported states: - `queued`: documents waiting for indexing. - `indexed`: documents already indexed. Supported sort values: - `created_at_asce` - `created_at_desc` - `updated_at_asce` - `updated_at_desc` - `indexed_at_asce` - `indexed_at_desc` [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Browse%20Documents) --- Source: https://docs.aivax.net/docs/rag/media-injector.html # Media Injector Media Injector turns a source file into focused, self-contained documents inside an AIVAX RAG collection. It examines the source, identifies materially useful knowledge, writes concise factual documents in the source's predominant language, and queues those documents for semantic indexing. Use Media Injector when you have a file whose useful knowledge has not already been split into retrieval-ready text. If you already have clean document strings, use [Create or Update Document or JSONL import](https://docs.aivax.net/docs/rag/collections.md#document-fields) instead; those paths are more predictable and avoid the additional processing needed to interpret a source file. ## When to use it Media Injector is useful for: - PDFs such as reports, manuals, and policies that contain several independent topics. - Images or scanned pages whose visible content should become searchable knowledge. - Audio and video whose material facts should be available through RAG. It is not a general file-storage feature and does not preserve the source as one searchable document. The output is a set of generated RAG documents. Review those documents after processing when wording, coverage, legal fidelity, or sensitive-data handling is important. Use direct document import when you need exact source wording, deterministic boundaries, stable document names, or application-controlled metadata. Use [Text Segmentation](https://docs.aivax.net/docs/rag/text-segmentation.md) when you only need cohesive source-text segments returned to your application without creating collection documents. Use this [RAG responsibility checklist](https://aivax.net/blog/a-vector-database-is-not-a-rag-system/) to decide which preparation steps to manage. ## How ingestion works In the AIVAX dashboard: 1. Open the target collection and choose **Import from files**. 2. Select one or more source files. 3. Optionally provide processing context. The same context is applied to every selected file. 4. Confirm the import. The dashboard uploads the files sequentially, and a separate job is created for each file after all of its chunks have reached AIVAX. 5. Follow the jobs under **Batch > Media Processing**. 6. After each job completes, review the generated documents and wait for their indexing state before testing [Semantic Search](https://docs.aivax.net/docs/rag/semantic-search.md). A job can be `queued`, `processing`, `completed`, `failed`, or `cancelled`. The dashboard reports the source file, elapsed time, number of documents produced, and current cost. Failed or cancelled jobs may be retried when their recoverable uploaded data is still available. Audio and video can be divided into time-based segments for processing. Segmentation is automatic and does not change the original file name shown for the job. See [Plans and limits](https://docs.aivax.net/docs/limits.md) for current upload limits. ## Define processing context Processing context is an optional instruction that helps Media Injector decide which facts are most valuable for your knowledge base. It is considered together with the source, but it is not treated as a factual source and cannot add facts that are absent from the file. Good context describes: - The source's identity and purpose. - The audience that will search the collection. - The topics, products, jurisdictions, periods, or fields that matter. - Ambiguous labels or internal terminology that the source itself establishes. - Content that should be deprioritized, such as repeated headers or administrative boilerplate. Example: ```text This is the 2026 support policy for Acme Cloud customers in Brazil. Prioritize eligibility rules, deadlines, plan differences, exceptions, and the steps a support agent must communicate. Ignore repeated page headers and signature blocks. ``` Avoid asking the mechanism to infer conclusions, supply missing information, or use outside knowledge. For example, do not instruct it to decide whether a contract is legally enforceable or to calculate values the source does not report. Processing context is different from a collection's context. Processing context guides this import only. Collection context describes the knowledge base to an AI Gateway when the collection is used later. Put source-specific ingestion guidance here; keep durable collection-wide guidance in the collection settings. ## Accepted source types The dashboard accepts four source groups for Media Injector: | Source group | Processing behavior | | --- | --- | | PDF documents | Reads document structure, text, and relevant visual content. | | Images | Interprets visible text and content to produce textual RAG documents. | | Audio | Interprets spoken and other relevant audible content; large files are segmented automatically. | | Video | Interprets relevant visual and audible content; large files are segmented automatically. | Use the original, accurate file extension because AIVAX uses it to identify the media type. A mislabeled extension can select the wrong processing path or make the job fail. Container and codec support can vary; if an audio or video file fails, convert it to a common format and retry. ## Generated documents Each generated item is designed to be a useful knowledge unit rather than a page-by-page transcription. Media Injector: - Prioritizes source identity, scope, principal facts, relationships, exceptions, and material distinctions. - Combines closely related facts instead of creating one document per label, table cell, or repeated value. - Ignores decorative text, pagination, repeated summaries, and incidental metadata unless they change meaning. - Preserves the source's language, terminology, displayed dates, and number formats. - Stops when the source has no materially new knowledge to add. Generated documents are tagged so they can be identified as automatically produced content. They are then indexed like other collection documents and incur the collection's normal text-embedding cost in addition to Media Injector processing. For retrieval-quality guidance after ingestion, see [Best Practices for RAG](https://docs.aivax.net/docs/rag/best-practices.md). In particular, inspect documents generated from tables, scans, and sources with repeated layouts before relying on them in production. ## Usage, pricing, and limits Media Injector usage depends on the source, optional context, generated questions and answers, cache reuse, and media tokens when applicable. Billing aggregates input, cached input, output, and media usage for the processing job without exposing the underlying processing model. See [Pricing](https://docs.aivax.net/docs/pricing.md#media-injector) for the final rates. Media Injector availability and operational limits depend on the account configuration. See [Pricing](https://docs.aivax.net/docs/pricing.md) and [Plans and limits](https://docs.aivax.net/docs/limits.md) before uploading files in production. --- Source: https://docs.aivax.net/docs/rag/semantic-search.html # Semantic Search The semantic search API searches one or more collections and returns the most relevant indexed documents for the supplied search terms. If your application already owns the candidate document strings, consider [Reflex](https://docs.aivax.net/docs/rag/reflex.md): a collection-less RAG search that ranks supplied documents without indexing or storage. Use managed semantic search when AIVAX should store and search a persistent corpus or when the corpus is too large to submit as candidates with every request. Compare [vector search with the complete RAG pipeline](https://aivax.net/blog/a-vector-database-is-not-a-rag-system/) before deciding what to build. After creating a collection, search it with complete terms that reflect the question a user would ask. The response can include the matched documents and their associated collection data for use in your application or AI Gateway flow. For the supported request, response, authentication, and error contract, use the API Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Semantic%20search) ## Reranking A reranker can adjust the order of candidates returned by semantic search. It does not search additional documents or recover text that the retrieval stage did not select. See [Rerankers](https://docs.aivax.net/docs/rag/reranking.md) for selection guidance. ## Multiple Terms Multiple terms cover alternative retrieval paths rather than requiring every term to match the same document. Use them for synonyms, alternative phrasings, or several acceptable ways to find an answer. If the user intent is one composite idea, send that idea as one complete term. For example, prefer: ```text How do I cancel an annual subscription without a penalty? ``` Over disconnected keywords: ```text cancellation annual subscription penalty ``` ## Search Quality A complete query usually performs better than a list of disconnected keywords because it preserves the relationship between concepts. If search returns poor results: 1. Confirm that the documents are indexed. 2. Query the collection directly before testing through an AI Gateway. 3. Compare complete questions with alternative phrasings. 4. Check whether the relevant document is too short, too long, or not self-contained. 5. Check whether the query language matches the document language. 6. If the gateway rewrites questions before searching, test with the plain query path to isolate rewriting issues. ## Collections MCP To expose AIVAX collections as tools for an external MCP client, see [Collections MCP](https://docs.aivax.net/docs/mcp-utilities/collections-mcp.md). For current service availability and account limits, see [Plans and Limits](https://docs.aivax.net/docs/limits.md). --- Source: https://docs.aivax.net/docs/rag/text-segmentation.html # Text Segmentation Text segmentation divides source documents into semantically cohesive strings that can be embedded or indexed in a RAG collection. It returns segments to your application; it does not create embeddings or store submitted documents. Use it when a source document needs reviewable, retrieval-ready boundaries before you create or update collection documents. If you already have focused, self-contained text, you can import it directly. If you want AIVAX to process source files into collection documents, see [Media Injector](https://docs.aivax.net/docs/rag/media-injector.md). See [where segmentation fits in a RAG pipeline](https://aivax.net/blog/a-vector-database-is-not-a-rag-system/), alongside storage, updates, and retrieval. ## Skip segmentation when the text is already focused Segmentation earns its keep on long, multi-topic sources — manuals, articles, transcripts — where one embedding per page would blur distinct subjects together. Do not segment text that is already one idea per unit: FAQ answers, product descriptions, short policies, or pre-chunked passages can go straight into the collection. Each unnecessary split adds indexing work and risks separating statements that only make sense together. ## What makes a good segment A good segment is the smallest span that still answers a question on its own: a complete statement or a tight group of statements about one subtopic, typically a paragraph or a short section. Segments are most useful when the source has clear headings, paragraphs, and complete statements — the segmenter preserves those boundaries instead of cutting mid-thought. Watch the failure modes of messy sources. Tables lose their headers, OCR drops line structure, transcripts ramble across topics, and exported documents repeat running headers on every page. Review results from those sources before indexing, and prefer cleaning the source (fixing headings, removing boilerplate) over asking the segmenter to guess around it. ## Prepare source text Supply complete source text whenever possible. Segments are most useful when the source has clear headings, paragraphs, and complete statements. Review results from tables, OCR, transcripts, or documents with repeated headers before indexing them. Use sanitization only when omitted content is genuinely irrelevant to retrieval. When exact source wording, legal fidelity, or full traceability matters, retain and review the source text instead. For the supported request, response, authentication, and error contract, use the API Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Segment%20text) For current service availability and account limits, see [Plans and Limits](https://docs.aivax.net/docs/limits.md). --- Source: https://docs.aivax.net/docs/rag/classification.html # Text Classification Use text classification to rank a fixed set of labels for one or more documents without training a custom classifier. AIVAX embeds every document and label with the default embedding model, compares their vectors using cosine similarity, and returns every label from the most similar to the least similar for each document. Before calling this endpoint, [create an API key](https://docs.aivax.net/docs/authentication.md) and make sure the account has a positive balance. ## Endpoint
POST /api/v1/generations/classify
## Request behavior | Property | Type | Required | Description | | --- | --- | --- | --- | | `documents` | `string[]` | Yes | One or more non-empty documents to classify. Results preserve this order and each document's zero-based index. | | `labels` | `string[]` | Yes | One or more non-empty labels. Every label is scored for every document. | Duplicate documents and labels are preserved. The endpoint always uses the current default embedding model and does not accept a model parameter, score threshold, or result limit. Example request: ```json { "documents": [ "Compare the total cost of two loans for a principal of $10,000 over 5 years at different annual rates", "Erklären Sie die Unterschiede zwischen Merge-Sort und Quicksort-Algorithmen in Bezug auf Zeitkomplexität, Platzkomplexität und Leistung in der Praxis.", "Write a poem about the beauty of nature and its healing power on the human soul" ], "labels": [ "Creative writing", "Complex problem", "Simple task" ] } ``` ## Read the response `results` contains one item for each input document. Each `scores` array contains every supplied label, ordered by descending cosine similarity. Labels with equal scores preserve their original order. ```json { "results": [ { "index": 0, "document": "Compare the total cost of two loans for a principal of $10,000 over 5 years at different annual rates", "scores": [ { "label": "Complex problem", "score": 0.98828 }, { "label": "Simple task", "score": 0.45272 }, { "label": "Creative writing", "score": 0.06823 } ] } ] } ``` A score measures vector similarity, not a calibrated probability. Compare scores within the same request and embedding model rather than interpreting a value as a percentage confidence. Negative scores are valid and remain in the response because the endpoint does not filter labels. Embedding usage is billed for text that requires inference and is associated with the authenticated API key. Repeated text may be served from an internal cache, reducing latency and costs. Because the endpoint returns every document-label pair, response size and comparison work grow with `documents × labels`. The embedded API reference contains the server-maintained request, response, authentication, and error details: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Classify%20documents) --- Source: https://docs.aivax.net/docs/rag/reranking.html # Rerankers Rerankers reorder an existing set of candidate documents for a query. They do not search a collection or recover text that is absent from the input. Use the autonomous reranking API when your application already owns the candidates, or use [Semantic Search](https://docs.aivax.net/docs/rag/semantic-search.md) to retrieve candidates from an AIVAX collection before reranking them. [Reflex](https://docs.aivax.net/docs/rag/reflex.md) is the collection-less search experience built on this same endpoint with the default ranker — see that page when you want retrieval-style ranking without managing a collection. ## Rerank documents directly Authenticate with an AIVAX API key and send a query with the candidate document strings. The API returns the candidates in relevance order, with the input position needed to associate each result with your application data. Use direct reranking when candidates are dynamic, come from another search system, or do not need to be stored in an AIVAX collection. Use managed semantic search when AIVAX should retrieve candidates from a persistent corpus. For the supported request, response, authentication, and error contract, use the API Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Rerank%20documents) ## Build candidates worth ranking Reranking only reorders what it receives, so candidate quality decides the ceiling. Keep each candidate string focused on one idea — a paragraph or a short section rather than a whole page — so the relevance score reflects one topic instead of an average over many. When candidates come from chunking, prefer boundaries that preserve complete statements; see [Text segmentation](https://docs.aivax.net/docs/rag/text-segmentation.md). Use the [RAG pipeline checklist](https://aivax.net/blog/a-vector-database-is-not-a-rag-system/) to distinguish preparation, retrieval, and ordering failures. Send enough candidates to cover plausible answers (over-retrieval first, precise ranking second) and use `top_n` to keep only the head of the ranked list. Use `min_score` to drop low-relevance tail results, but calibrate the threshold on your own queries: score scales differ between rankers and a cutoff tuned on one workload rarely transfers to another. ## Choosing a reranker Start with the default reranker unless you have a measured reason to select another available option. The default is Reflex; alternatives include lexical matching, reciprocal-rank fusion of several signals, and third-party cross-encoders. Note that fusion (`rrf`) is retrieval-only and is rejected by this endpoint — it exists for combining signals inside collection search, not for autonomous ranking. To compare options, fix a set of representative queries with known relevant documents from your own workload — covering your languages, document lengths, and jargon — and measure whether swapping rankers moves the right document upward. If the relevant document is missing from the candidates entirely, improve candidate retrieval, chunking, query formulation, or candidate count before comparing rerankers: no ranker recovers what was never submitted. For current availability, supported options, and account limits, see the API Reference and [Plans and Limits](https://docs.aivax.net/docs/limits.md). --- Source: https://docs.aivax.net/docs/rag/reflex.html # Reflex Reflex is AIVAX's collection-less search for RAG. Send a query together with candidate document strings and receive the most relevant items in ranked order—without indexing, storing, or maintaining a RAG collection first. Reflex is the default ranker of the autonomous [reranking endpoint](https://docs.aivax.net/docs/rag/reranking.md): calling that endpoint without a `model` selects Reflex. This page covers when to reach for Reflex; that page covers candidate preparation and ranker comparison in depth. Use Reflex when your application already owns the candidate documents, the candidate set changes frequently, or you want a retrieval step without collection indexing and storage. Use [Semantic Search](https://docs.aivax.net/docs/rag/semantic-search.md) when AIVAX should store, index, and search a persistent knowledge base or narrow a corpus that is too large to submit as candidates on every request. ## Reflex or Semantic Search? | Choose Reflex when... | Choose Semantic Search when... | | --- | --- | | Your application already has the candidate document strings. | Documents should live in managed AIVAX collections. | | You need retrieval immediately, without an indexing step. | The knowledge base is persistent and searched repeatedly. | | The candidate set is dynamic or request-specific. | The corpus is too large to submit as candidates on every request. | | You want collection-less ranking. | You want collection filtering, stored metadata, document references, and managed retrieval. | Reflex returns ranked text candidates; it does not generate an answer. Pass the selected documents to your language model or AI Gateway as RAG context. ## Use Reflex Call the reranking API with a query and the candidate documents your application wants to compare. Results are returned in relevance order and retain the input position needed to associate them with your application data. Use concise, focused candidate documents. Reranking can improve their order, but it cannot recover information that was not included in the candidates. If the expected document is consistently absent, improve candidate selection, chunking, or query formulation before tuning the ranker. Reflex accepts up to 10,000 candidate documents per request and returns at most 200 ranked results. If your candidate pool exceeds 10,000 documents, narrow it first — with lexical pre-filtering, a cheap first-pass rank, or collection retrieval — and let Reflex order the shortlist. For the supported request, response, authentication, and error contract, use the API Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Rerank%20documents) ## Use Reflex with RAG Reflex is also the default reranker after AIVAX retrieves candidates from RAG collections. In this flow, it can improve the order of retrieved candidates but cannot recover a document that the retrieval stage did not select. If relevant documents are consistently absent, adjust retrieval, chunking, query formulation, or candidate count before tuning reranking. Free, Pro, and Max include a daily reranking allowance for Reflex, shared between autonomous calls and RAG reranking. Cached and uncached input both consume this allowance. It is separate from the RAG embedding allowance and the processing-time cap; other rerankers are billed normally. For relative plan capacity and coverage rules, see [Plans and Limits](https://docs.aivax.net/docs/limits.md#included-daily-subscription-allowances). --- Source: https://docs.aivax.net/docs/rag/best-practices.html # Best Practices for RAG Use this guide when preparing documents for AIVAX collections and semantic search. The search pipeline indexes document text as embeddings, stores the resulting vectors, and later compares user queries with those indexed vectors. Search quality depends heavily on how clear, focused, and self-contained each document is. ## Document Size A document should represent one limited piece of knowledge. As a practical target, keep most documents between 20 and 700 words. This range is not a hard API rule, but it usually gives the embedding model enough context without mixing unrelated topics. The current indexer also records warnings for unusually small or large documents: - Documents below about 10 tokens are accepted, but may be too small to retrieve reliably. - Documents above about 1,562 tokens are accepted, but the indexer truncates the text used for embedding to about 5,000 characters and records a warning. If a document is longer than that, split it before indexing. Use paragraphs, sections, policy clauses, FAQ entries, product descriptions, or other logical units. ## What to Avoid - Empty, tiny, or title-only documents. - Entire PDFs, chapters, logs, or manuals as one document. - Multiple unrelated subjects in one document. - Text that depends on surrounding pages to make sense. - Mixed languages inside the same document, unless the user is expected to search that way. - Raw JSON, code, tables, or logs without a short natural-language explanation. - Generic pronouns and references such as "it", "this process", or "the product" when the document does not identify the subject. ## What to Do - Give each document a clear name and focused text. - Put the subject near the start of the document. - Use natural language similar to the way users ask questions. - Repeat important identifiers, product names, policy names, acronyms, and terms when they matter. - Keep one document focused on one answerable topic. - Use `__tags` to organize documents operationally. - Use `__ref` to group chunks that belong to the same logical source. - Use `__meta` for structured data your application needs to keep, such as source URL, version, author, or publication date. Example: Prefer: ```text The color of the 2015 Honda Civic registered in fleet record CAR-123 is yellow. ``` Avoid: ```text The car is yellow. ``` The first version can be retrieved and understood without external context. ## Metadata, Tags, and References Only the document text is embedded for semantic matching. Metadata is returned with results and can be useful for applications, audits, source links, versioning, or display, but it should not replace searchable text. Use tags for maintenance and filtering, not as the only place where important meaning appears. If a user may search for "refund policy", those words should appear in the document text, not only in a tag. Use references when several chunks represent the same source item. When reference expansion is enabled, a matched chunk can return other documents that share the same reference. ## Chunking Larger Sources When importing PDFs, spreadsheets, web pages, or manuals, inspect the generated chunks before relying on the collection. Remove repeated headers, footers, navigation menus, broken tables, boilerplate, and irrelevant disclaimers when possible. See [what a vector database leaves to the application](https://aivax.net/blog/a-vector-database-is-not-a-rag-system/) before choosing an ingestion workflow. Good chunks usually include: - A title or source heading. - The immediate section context. - The complete rule, answer, instruction, or explanation. - Enough surrounding text to answer a question without neighboring pages. Poor chunks often contain: - Half of a table row. - A sentence that depends on the previous page. - Several policies mixed into one block. - Repeated layout text from the original file. ## Query Quality Semantic search works best when documents and queries use compatible language. If users ask full questions, prepare documents that contain complete explanations. If users search by product codes, policy IDs, plan names, or procedure names, include those identifiers in the text. When search results are poor, check the basics first: - Confirm the documents are indexed. - Search the collection directly before testing through an AI gateway. - Try a complete question instead of isolated keywords. - Compare the query language with the document language. - Review whether the relevant answer is split across too many small chunks or buried inside a very large one. Well-prepared documents make RAG predictable: the model receives clearer source material, search returns fewer irrelevant matches, and answers become easier to audit. --- Source: https://docs.aivax.net/docs/inference/ai-gateway.html # AI Gateway An AI Gateway is a persistent inference configuration. It lets you call a gateway by model name while AIVAX applies the gateway's model settings, instructions, RAG collections, tools, skills, workers, moderation, and context controls. Use a gateway when the same behavior must be reused by multiple clients or changed without redeploying the calling application. ## How to think about a gateway A direct call to `/v1/chat/completions` can call an integrated AIVAX model directly, for example `@openai/gpt-5-mini`. A gateway stores the decisions you do not want to repeat on every request: - Model provider and model name. - System instructions, remote instruction sources, user prompt template, and assistant prefill. - RAG collections, result limits, score threshold, reranker, reference behavior, and query strategy. - OpenAI-compatible tools, AIVAX built-in tools, MCP tools, protocol functions, skills, and the optional bash environment. - Context window behavior, tool-message truncation, moderation, workers, model routing, and tool-call handling. This creates a responsibility boundary. The client application sends messages and optional request overrides. The gateway administrator controls the operational policy. In production, start with a conservative configuration: clear system instructions, a model that supports the required modalities and tools, one well-prepared RAG collection, and only the tools that are actually needed. Adding many tools, skills, or collections increases input tokens, cost, and the chance of the model choosing the wrong path. ## Models and gateway names There are three common ways to choose what `/v1/chat/completions` uses: - Use an integrated AIVAX model tag, usually beginning with `@`. - Use a gateway full ID. - Use a gateway slug in the format `name:final-id`, such as `support:50c3`. Private API keys can resolve a gateway by full ID or by slug. Public API keys are more restricted: they can use AI gateways only by full ID, and only a limited set of chat completion request parameters is accepted. When choosing a model, validate three points before putting it into production: - The model supports the input modalities you intend to send, such as image, audio, video, or file. - The model supports function calling if the gateway uses tools, RAG through `QueryFunction`, MCP, protocol functions, skills, or built-in functions. - The model accepts the parameters you configure. Some integrated models reject assistant prefill, temperature, stop sequences, or reasoning effort. Gateways can also use model routing. For the complexity router, AIVAX classifies the latest user request as low, medium, or high complexity, selects the configured model for that level, and emits `X-Model-Routed-Complexity` on the HTTP response when available. ## Using an AI Gateway AIVAX provides an OpenAI-compatible chat completions endpoint: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Inference%20(chat%20completions)) Gateway values can be overridden by the request for supported parameters such as `temperature`, `top_p`, `seed`, `reasoning_effort`, `max_completion_tokens`, `stop`, `tools`, `response_schema`, `response_format`, `builtin_tools`, `multimodal_preprocess`, and `tool_invocation_explanations`. For direct inference behavior, including response rendering options, see [Inference](https://docs.aivax.net/docs/inference/inference.md). ## Using SDKs Because the endpoint follows the OpenAI chat completions shape, you can use existing OpenAI-compatible SDKs. ```python from openai import OpenAI client = OpenAI( base_url="https://inference.aivax.net/v1", api_key="" ) response = client.chat.completions.create( model="my-gateway:50c3", messages=[ {"role": "user", "content": "Explain why AI gateways are useful."} ] ) print(response.choices[0].message.content) ``` OpenAI-compatible inference uses `/v1/chat/completions`. The `/v1/responses` endpoint is not supported. ## Recommended production configuration Write the system instructions so the model understands its role, audience, sources of truth, and limits. Include when to use RAG, when to use tools, and how to respond when information is unavailable. Avoid repeating operational settings that already exist in the gateway, such as truncation limits or tool lists. For RAG, link collections with short, self-contained, well-named documents. Choose the query strategy based on the conversation type: - `Plain`: Uses the last user message as the search term. - `Concatenate`: Joins the last configured number of user messages line by line. - `UserRewrite`: Rewrites recent user messages into one or more search queries using a resolver model. - `FullRewrite`: Rewrites recent user and assistant messages using a resolver model. - `QueryFunction`: Exposes a search function to the model instead of injecting a search result before inference. For tools, enable only those with a clear role. Built-in tools cover common capabilities such as current date and time, web search, opening URLs, code execution, image generation, document generation, page generation, calendar actions, memory, HTTP requests, and X post lookup. External MCP is better when you already have an MCP server with business tools. Protocol functions are useful when you want to expose specific HTTP callbacks to the model without installing a full MCP server. When enabling memory, define what the assistant may retain and how your application will review and remove stored records. The [memory-poisoning guide](https://aivax.net/blog/persistent-memory-is-a-write-path/) outlines these controls. Use a tool handler only when the selected model needs help producing tool calls. The available handler is `react.v1.selfcall`; `native` or no value uses the model's native tool calling. Use workers when an external system must decide something during the inference flow. A worker can block a message, rewrite context, add tools, or replace a server-side tool result. Because the worker is called in the critical path, keep it fast and deterministic. Validate your gateway configuration with a [simulated-user and LLM-judge scenario](https://aivax.net/blog/introducing-agentic-tests/) before relying on it in production. ## Inference MCP To expose an integrated model or AI Gateway as a tool for an external MCP client, see [Inference MCP](https://docs.aivax.net/docs/mcp-utilities/inference-mcp.md). --- Source: https://docs.aivax.net/docs/inference/inference.html # Inference AIVAX exposes an OpenAI-compatible `chat/completions` API with additional AIVAX parameters. The additions are optional and are designed to support gateways, RAG, built-in tools, structured responses, multimodal pre-processing, model routing, and billing metadata. Use this page for direct inference calls. Use [AI Gateway](https://docs.aivax.net/docs/inference/ai-gateway.md) when the same configuration must be reused or centrally managed. ## Endpoint
POST /v1/chat/completions
The endpoint also has the API alias `/api/v1/chat/completions`. Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Inference%20(chat%20completions)) ## Provider routing Some integrated models are available through more than one provider. Provider routing lets AIVAX choose among those providers without changing the model requested by your application. This differs from model routing, which can select a different model based on request complexity. AIVAX considers providers that are currently available and compatible with the request. If only one provider is eligible, the routing preference does not change the result. Provider routing applies only to integrated AIVAX models; a bring-your-own-key gateway uses the provider endpoint configured in that gateway. Available routing preferences are: | Preference | Behavior | |---|---| | `Balanced` | Balances price, speed, and quality. This is the default. | | `Cheapest` | Selects the provider with the lowest applicable input and output token price. | | `Fastest` | Prioritizes the provider with the highest available throughput. | | `Quality` | Selects the provider that AIVAX ranks highest for quality, without optimizing for price or speed. | ### Configure routing in an AI Gateway Use an AI Gateway when the same routing preference should apply to every request. In the gateway editor, select an integrated model, open **Routing preference**, choose the preferred strategy, and save the gateway. The equivalent gateway configuration uses `parameters.routingOption`: ```json { "name": "Cost-optimized assistant", "parameters": { "baseAddress": "@integrated", "modelName": "YOUR_INTEGRATED_MODEL", "routingOption": "Cheapest" } } ``` After saving, call the gateway normally by using its ID or slug as `model`. AIVAX applies the stored routing preference while preserving the gateway's instructions, tools, RAG configuration, and other settings. See [AI Gateway](https://docs.aivax.net/docs/inference/ai-gateway.md) for the complete gateway workflow. ### Override routing in `chat/completions` Use `routing_preset` to choose a provider strategy for one request. The override works with a direct integrated model or an AI Gateway that uses an integrated model: ```json { "model": "YOUR_INTEGRATED_MODEL_OR_GATEWAY_ID", "messages": [ { "role": "user", "content": "Summarize this incident report." } ], "routing_preset": "Fastest" } ``` The accepted values are `Balanced`, `Cheapest`, `Fastest`, and `Quality`. The request value overrides the gateway's saved `routingOption` for that request only; it does not update the gateway. Because `routing_preset` is an AIVAX extension, send it as an extra request-body field when using an OpenAI-compatible SDK. Request-level routing overrides require a private API key. ## Input and multimodality AIVAX accepts OpenAI-compatible message content parts for text, images, audio, videos, and files. The selected model must support the modality unless you ask AIVAX to pre-process the media into text. ```json { "model": "@google/gemini-3-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe these inputs briefly." }, { "type": "image_url", "image_url": { "url": "data:image/png;base64,", "detail": "auto" } }, { "type": "input_audio", "input_audio": { "data": "base64-encoded-audio", "format": "wav" } }, { "type": "file", "file": { "filename": "document.pdf", "file_data": "data:application/pdf;base64," } } ] } ] } ``` Supported content part mappings: - `text`: Plain text. - `image_url`: Image content. `image_url.url` can be an external URL or a base64 data URL. `image_url.detail` can be `low`, `high`, or `auto` when the model supports it. - `video_url`: Video content. `video_url.url` can be an external URL or a base64 data URL. Prefer URLs for large videos. - `input_audio`: Audio content. `input_audio.data` is base64 audio data, and `input_audio.format` names the format. - `file`: File content. `file.filename` names the file, and `file.file_data` can be an external URL or a base64 data URL. For video input, send a `video_url` content part. The following example uses a base64 Data URL; prefer a publicly reachable URL for large videos: ```json { "model": "@google/gemini-3-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Summarize the main actions in this video and identify any visible safety risks." }, { "type": "video_url", "video_url": { "url": "data:video/mp4;base64," } } ] } ] } ``` External links must be accessible to AIVAX without authentication, firewall restrictions, or JavaScript-only rendering. Failed downloads, redirects, blocked URLs, unsupported formats, or provider-specific size limits can fail the inference. You can also send a simple text request with `prompt`: ```json { "model": "@google/gemini-3-flash", "prompt": "Say hello" } ``` ## Request idempotency Set `idempotency_key` when your integration needs repeat calls to update the same stored conversation record instead of creating a new conversation token. AIVAX uses this value to correlate the AI Gateway context and conversation logging. ```json { "model": "your-model-or-gateway-id", "messages": [ { "role": "user", "content": "Summarize order 123." } ], "idempotency_key": "order-123-summary" } ``` The value must be a non-empty string with 128 characters or less. When omitted, AIVAX generates a conversation token automatically. ## Request metadata Set `metadata` to attach string key/value information to the inference request. AIVAX stores this object with the logged conversation and exposes it to gateway events, so it is useful for operational correlation such as an order ID, tenant, workflow, or internal trace key. ```json { "model": "your-model-or-gateway-id", "messages": [ { "role": "user", "content": "Summarize this support ticket." } ], "metadata": { "ticket_id": "SUP-1042", "workflow": "support-triage" } } ``` `metadata` must be a JSON object whose property names and values are strings. Do not place secrets, credentials, payment data, or large payloads in this field. ## Response and conversation records The standard `/v1/chat/completions` response envelope includes `generation_context`. Its `generated_usage` entries contain `sku`, `amount`, `unit_price`, `quantity`, and `description`. Set `json_only: true` to return only the final JSON without this envelope. When conversation logging is enabled, the stored record includes its ID, origin, model name, request ID, response schema, tools and tool input schemas, usage, linked resources, created and updated timestamps, token count, external user ID, error message, messages, and metadata. Gateway and API key context is available through the linked resources. Use `idempotency_key` and `metadata` to correlate these records with your own workflow. ## Multimodal pre-processing Use `multimodal_preprocess` when the main model should receive a textual description of media instead of the original media object. This is useful for text-first models or when you want AIVAX to normalize files before the main inference. ```json { "model": "@metaai/llama-3.3-70b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this file briefly." }, { "type": "file", "file": { "filename": "document.pdf", "file_data": "data:application/pdf;base64,BASE64_PDF_CONTENT" } } ] } ], "multimodal_preprocess": "File" } ``` Available pre-processing flags are: - `Image` - `Audio` - `Video` - `File` - `OtherFiles` - `All` The resolver caches media descriptions by content hash for reuse. `Image`, `Audio`, `Video`, and PDF `File` pre-processing use auxiliary multimodal inference. Supported non-PDF files use local text extraction. Multimodal inputs can have account requirements. Review [Pricing](https://docs.aivax.net/docs/pricing.md) and [Plans and limits](https://docs.aivax.net/docs/limits.md) before using them in production. When a multimodal inference fails, narrow down the problem: 1. Test a simple text message with the same model. 2. Test one small attachment. 3. Test the same attachment with `multimodal_preprocess`. 4. Review the URL, format, size, and model modality support. ## Structured responses AIVAX supports structured responses through `response_schema`, `response_format`, and `json_only`. ```json { "model": "@google/gemini-2.5-flash", "prompt": "Search for recent news about electric vehicles.", "stream": true, "builtin_tools": { "tools": [ "WebSearch" ], "options": { "web_search_mode": "full" } }, "response_schema": { "type": "object", "properties": { "news": { "type": "array", "items": { "type": "object", "properties": { "title": { "type": "string", "description": "News title" }, "summary": { "type": "string", "description": "News summary" } }, "required": ["title", "summary"] } } }, "required": ["news"] } } ``` `response_schema` enables JSON Healing. AIVAX asks the model for JSON, extracts JSON from the generated text or markdown blocks, validates it against the schema, and retries with validation feedback until the output is valid or the attempt limit is reached. Read more on [Structured responses](https://docs.aivax.net/docs/inference/structured-responses.md). If your application cannot parse or validate the result, follow the [invalid JSON troubleshooting guide](https://aivax.net/blog/structured-output-healing-boundary/) before increasing the retry budget. ## On-demand functions Use `builtin_tools` to enable AIVAX built-in tools for a direct request without creating a gateway: ```json { "model": "@google/gemini-2.5-flash", "prompt": "Search for recent news about electric vehicles.", "stream": true, "builtin_tools": { "tools": [ "WebSearch" ], "options": { "web_search_mode": "full", "web_search_max_results": 5 } } } ``` Built-in tools include `DateTime`, `WebSearch`, `AdvancedWebUsage` (disabled; returns an unavailable response; see [Changelogs](https://docs.aivax.net/docs/changelogs.md)), `OpenUrl`, `Code`, `Request`, `Calendar`, `Remember`, `GenerateWebPage`, `GenerateDocument`, `XPostsSearch`, and `ImageGeneration`. `DateTime` exposes `get_date_time`, a no-argument tool returning the current date, time, English weekday, time zone, UTC offset, and ISO 8601 timestamp. Set `builtin_tools.options.dateTimeTimeZone` to an IANA identifier; the default is `America/Los_Angeles` (Pacific Time), with automatic daylight-saving adjustments. This setting is independent of the user's browser time zone. See [Current Date and Time](https://docs.aivax.net/docs/tools/builtin-tools.md#current-date-and-time) for configuration and output examples. On-demand tools are suitable for occasional calls, prototypes, and integrations that do not need a persistent gateway. If the same application always uses the same tools, prefer configuring them in an AI Gateway so the policy is centralized. ## Custom provider request body When a gateway uses a provided API key and an OpenAI-compatible provider endpoint, `extra_body` can merge custom JSON into the provider request body: ```json { "model": "my-custom-model:abc4", "messages": [ { "role": "user", "content": "Explain the tradeoff." } ], "extra_body": { "reasoning": { "enabled": true } } } ``` `extra_body` is not allowed with integrated AIVAX models. ## Tool explanations Set `tool_invocation_explanations: true` to ask AIVAX to include explanation fields in server-side tool arguments. When the model supplies `_tool_reason` and `_tool_goal`, `servertool.explanation` contains a client-friendly copy: ```json { "model": "@x-ai/grok-4.3", "messages": [ { "role": "user", "content": "What's the weather forecast for today?" } ], "stream": true, "builtin_tools": { "tools": ["WebSearch"] }, "tool_invocation_explanations": true } ``` Example stream event: ```json { "choices": [], "servertool": { "name": "web_search", "id": "call-example-id-0", "contents": "{\"query\":\"weather forecast today\",\"_tool_reason\":\"Searching for today's weather forecast online\",\"_tool_goal\":\"I need current weather information to answer accurately.\"}", "state": "Created", "explanation": { "reason": "Searching for today's weather forecast online", "goal": "I need current weather information to answer accurately." } }, "usage": null } ``` ## Response rendering mode Set `rendering_mode: "textual_blocks"` when your client wants AIVAX to place reasoning and server-side tool activity in the same textual response flow that the chat UI already renders. This is useful for clients that build a single response timeline and want to turn reasoning and tool activity into visible components without keeping separate event-handling paths for every marker type. ```json { "model": "@openai/gpt-5-mini", "messages": [ { "role": "user", "content": "Search for recent product updates and summarize the important changes." } ], "stream": true, "builtin_tools": { "tools": ["WebSearch"] }, "rendering_mode": "textual_blocks" } ``` In this mode, reasoning can be emitted as `` and `` blocks, assistant-facing text can be emitted as `` blocks, and server-side tool markers can appear as tool result elements such as `
`. Treat these blocks as presentation markers inside the response stream: parse them into chat timeline components, collapsible reasoning sections, assistant answer fragments, or tool status rows, but do not blindly concatenate every marker into the final assistant answer. Clients that do not understand this markup should keep the default rendering mode and handle the structured stream events directly. In the default mode, reasoning arrives through `delta.reasoning`, and server-side tool activity arrives through `servertool` events. Preserve the order in which stream events arrive so reasoning, tool activity, partial content, and the final answer remain in the same response timeline. ### Raw multi-turn example The example below shows the shape of a streamed response when server-side reasoning is visible to the client, `tool_invocation_explanations` is enabled, and `textual_blocks` is used to keep the response timeline textual. The exact tool result attributes can vary by renderer, but the important behavior is the ordering: reasoning, assistant answer fragments, tool activity, more reasoning, and the final answer can all belong to the same assistant turn. ```json { "model": "my-custom-model:abc4", "messages": [ { "role": "user", "content": "Which cheap and fast multimodal models should I use for security camera analysis?" } ], "stream": true, "builtin_tools": { "tools": ["WebSearch"] }, "tool_invocation_explanations": true, "rendering_mode": "textual_blocks", "extra_body": { "reasoning": { "enabled": true } } } ``` Raw streamed assistant timeline: ```text The user is asking for cheap, fast multimodal models for security camera analysis. I should list available AIVAX models and search the documentation before recommending options. I will check the available multimodal models and identify the best options for security camera analysis.
aivax_list_modelsListing the available models in AIVAX
aivax_search_contextSearching documentation about multimodal models and image analysis in AIVAX
The relevant models should support VideoInput or ImageInput, have low input cost, and be fast enough for camera workflows. I found several candidates and should rank them by cost, speed, and modality support.
For security camera analysis, prioritize models with VideoInput, low input pricing, and high speed. Model availability and prices change over time; the picks below are example output — see [Pricing](https://docs.aivax.net/docs/pricing.md) for current values. Top picks: 1. @google/gemini-2.5-flash-lite: fast, inexpensive, and supports video. 2. @qwen/qwen3.5-9b: low input cost in this example output with video support. 3. @amazon/nova-lite: low input cost and a large context window. Use VideoInput for clips when possible. If a model only supports ImageInput, extract frames from the camera stream before sending them. ``` When the user replies, keep the conversation history focused on the user-visible assistant result. Store reasoning and tool details as timeline or audit metadata if your product needs them, but do not turn them into a new user message. The assistant message should use the content from the final `` block, not the full reasoning transcript. ```json { "model": "my-custom-model:abc4", "messages": [ { "role": "user", "content": "Which cheap and fast multimodal models should I use for security camera analysis?" }, { "role": "assistant", "content": "For security camera analysis, prioritize models with VideoInput, low input pricing, and high speed.\n\nModel availability and prices change over time; the picks below are example output — see [Pricing](https://docs.aivax.net/docs/pricing.md) for current values.\n\nTop picks:\n\n1. @google/gemini-2.5-flash-lite: fast, inexpensive, and supports video.\n2. @qwen/qwen3.5-9b: low input cost in this example output with video support.\n3. @amazon/nova-lite: low input cost and a large context window.\n\nUse VideoInput for clips when possible. If a model only supports ImageInput, extract frames from the camera stream before sending them." }, { "role": "user", "content": "Now recommend one model for real-time alerts and one for deeper review." } ], "stream": true, "builtin_tools": { "tools": ["WebSearch"] }, "tool_invocation_explanations": true, "rendering_mode": "textual_blocks", "extra_body": { "reasoning": { "enabled": true } } } ``` ### Presentation guidance During generation, reasoning is useful because it lets the user follow what the model is doing before the final answer exists. The assistant can "speak" while it reasons by emitting user-facing process updates or provisional answer fragments. These updates may be interleaved with reasoning blocks, tool calls, and partial answer content as the response develops. Once the final assistant answer is generated, that answer becomes the main product of the inference. The intermediate reasoning is still useful for audit, orientation, and debugging, but it usually stops being the user's primary goal. Collapse or minimize reasoning by default after completion so the final answer receives the strongest visual emphasis, while keeping the process available for users who want to inspect it. Use progressive disclosure around that lifecycle. Reasoning can be visible while the model is still working, then become a quieter, secondary element after the final answer appears. Tool activity should read as status, not speech: use concise labels such as "Searching", "Opening source", "Running tool", "Finished", or "Failed", and keep each tool invocation grouped as one timeline item even if its state changes over time. A good visual hierarchy is: - Assistant answer: highest prominence, normal reading typography, part of the main conversation. - In-progress reasoning: visible enough to show what the model is doing while the response is being generated. - Completed reasoning: lower prominence, subdued color or container, collapsed or minimized by default. - Tool blocks: compact status rows with clear loading, success, and error states. - Raw details: hidden by default unless the client is a developer, audit, or debugging surface. Avoid exposing noisy internals directly to end users. Show tool names, states, source labels, or short summaries when they help the user understand what happened. Hide raw arguments, large payloads, and implementation details unless the user explicitly asks for details or the product surface is built for technical inspection. For accessibility, make every collapsed block keyboard-toggleable, give each status row a readable label, avoid relying on color alone for state, and keep motion subtle. A streaming response should feel stable while it updates: new reasoning or tool rows can appear in order, but existing content should not jump around or force the user to lose their reading position. ## Direct call or gateway Use a direct call for simple tasks, tests, internal routines, and integrations where the application controls the model, prompt, tools, and context for each request. Use an AI Gateway when behavior needs to be stable, auditable, and reusable. Gateways are better for support assistants, chat bots, RAG agents, permanent tools, workers, skills, and configurations shared by multiple clients. --- Source: https://docs.aivax.net/docs/inference/agentic-tests.html # Agentic Tests Agentic Tests evaluates how an AI Gateway behaves across a complete, goal-oriented conversation instead of grading one isolated response. AIVAX simulates the user's next message, sends each turn to the selected gateway, and uses an independent judge to determine whether the conversation reached its goal, remains recoverable, or has persistently moved away from the expected outcome. Use Agentic Tests to create repeatable regression checks for support, sales, onboarding, tool use, RAG, and other multi-turn agent flows. Because a test runs through the configured AI Gateway, it exercises the gateway's model, instructions, tools, skills, knowledge, and inference settings together. ## Persistent tests in the dashboard Open **Agentic Tests** in the AIVAX dashboard to create and manage reusable test cases. A test stores: - the AI Gateway under test; - a goal that describes the expected conversational outcome and is shared with the simulated user and judge; - optional validation criteria used only by the judge; - optional starting messages, external resources, and an external user identifier; - simulated-user sampling, turn limits, exit behavior, and evaluation thresholds; - an optional recurring schedule; - failure and recovery notification settings. The test definition is reusable. Each execution creates a separate run, so changing a test later does not replace the history already collected for previous runs. ### Create a useful test Write the goal from the simulated user's perspective: describe who they are, what they want, and how they should progress through the conversation. Do not write it as instructions for the assistant. The goal is shared with both the simulated user, which pursues it, and the judge, which evaluates it. For example: > You are choosing a plan for your team. Explain your team's size and needs when asked, ask which plan fits, and continue until you understand the recommendation and how to sign up. Use **Validation criteria** for optional requirements that should affect only the judge's evaluation, not the simulated user's behavior. For example: > The recommendation must name the selected plan and connect it to the stated team size. The final response must include a direct signup step. Keeping these criteria separate prevents the simulated user from unnaturally steering the conversation toward the checks that the judge will apply. Use **Start messages** when the scenario requires an established context, such as a customer objection, a prior assistant response, or a specific point in an existing flow. Use `external_user_id` when the gateway behavior depends on an identity from your own application. The value is forwarded to gateway inference for every run of that test. Use `resources` to give the simulated user and judge shared context that does not belong in the initial conversation. Provide up to 16 objects with a `type` and non-empty `data` value. `Text` uses `data` as literal context; `RemoteResource` retrieves the content from the URL in `data`. For example: ```json { "resources": [ { "type": "Text", "data": "The customer has a 14-day refund window." }, { "type": "RemoteResource", "data": "https://example.com/refund-policy" } ] } ``` Resources are visible to the simulated user and judge; they are not sent to the gateway under test as conversation history. They do not add knowledge to the gateway. If the assistant must retrieve the same material, make it available through the gateway's own knowledge or tools. The simulated user may still reveal resource information naturally in its messages, so do not treat resources as hidden judge-only criteria. Remote content can change between test runs and contributes to usage. Use only trusted, publicly reachable URLs whose content is appropriate for the test. A focused test usually gives more actionable results than one broad scenario. Separate unrelated goals into different tests so that a failure identifies the behavior that regressed. Follow the [multi-turn conversation testing walkthrough](https://aivax.net/blog/introducing-agentic-tests/) to define observable goals, run a simulated conversation, and inspect recovery. ### Validation hooks Persisted Agentic Tests may call external validation hooks during a run. Configure the `hooks` array when creating or updating a test: ```json { "hooks": [ { "event": "before-test", "url": "https://validator.example/hooks/agentic-tests" }, { "event": "after-test", "url": "https://validator.example/hooks/agentic-tests" }, { "event": "before-inference", "url": "https://validator.example/hooks/gateway" }, { "event": "after-inference", "url": "https://validator.example/hooks/gateway" }, { "event": "context-changed", "url": "https://validator.example/hooks/agentic-tests" } ] } ``` The supported events are: | Event | When it is sent | Event data | |---|---|---| | `before-test` | Before the first simulated turn. | `gateway`, `goal`, and `metadata`. | | `after-test` | After the test reaches a normal terminal outcome and before the final event is emitted. | The final outcome, reason, state, turn number, score, conversation delta, and loss streak. | | `before-inference` | Once before the primary gateway inference of each turn. | `turn_number` and the current `messages` array. | | `after-inference` | Once after the primary gateway inference of each turn. | `turn_number` and the current `messages` array. | | `context-changed` | Once per turn after the gateway response has completed. | `turn_number` and the current `messages` array. | `before-inference` and `after-inference` are available only when the test targets an AI Gateway. Hooks are supported for persisted dashboard tests and scheduled runs; the direct SSE validation endpoint does not accept `hooks`. Each hook receives a worker-compatible JSON envelope: ```json { "testId": "", "runId": "", "gatewayId": "", "moment": "2026-08-16T03:00:00Z", "event": { "name": "before-inference", "data": { "turn_number": 1, "messages": [] } } } ``` The hook URL must be an absolute HTTP or HTTPS URL without embedded credentials and must not resolve to localhost, loopback, private, link-local, or other blocked local addresses. When the account has a hook key, AIVAX also sends `X-Request-Nonce`; validate it before trusting the payload. Hook requests use `POST` with `Content-Type: application/json`. Hook responses follow the worker convention: any `2xx` response continues the run; a non-`2xx` response or HTTP request failure interrupts it immediately. The response body does not select another action. The interrupted run is stored as `failed`, emits a terminal result with `reason: "validation_hook_interrupted"`, and includes hook call audit entries in the run `result.hooks` array. Audit entries contain the event, URL, timestamp, status or error, whether execution continued, and up to 4,000 characters of the response body. ### Run and inspect a test Select **Run test** to queue an execution. Runs can be `pending`, `running`, `succeeded`, `failed`, or `cancelled`. Both the rate of new runs and account-level concurrency depend on the current plan. See [Plans and limits](https://docs.aivax.net/docs/limits.md#plan-limits) for current values. Manual runs, scheduled runs, and direct API evaluations share one new-run quota across the account's API keys. A persisted run counts when it is queued and does not count again when execution begins. Conversation turns do not consume additional run units, although applicable inference limits still apply. A manual run request above the quota returns HTTP 429 without creating a run; wait for the rate-limit window to clear before retrying. A run processes its conversation sequentially, while eligible runs from the same account can execute concurrently. Every turn checks that the account can continue operating. A run can fail if the balance is exhausted or inference cannot continue, and a pending or running run can be cancelled from the dashboard. The run inspector retains: - simulated-user, assistant, and judge messages in chronological order; - a precise timestamp for every retained message; - prompt, cached prompt, and completion token usage per message; - every judge opinion, including its reasoning, score, state, and trajectory values; - the final evaluation result, failure information, and total cost charged to the run. Use the judge opinions to identify the turn where the conversation improved, became at risk, succeeded, or entered a persistent loss. The run detail can also be exported as JSON for offline review. ### Schedule recurring tests A test can run automatically from a standard five-field cron expression. The minimum supported interval is five minutes. For example, `*/15 * * * *` runs every 15 minutes. If the account's new-run quota is exhausted, a due scheduled test waits for a later scheduling check without creating a run. Scheduling does not bypass the quota or reserve capacity separately from manual runs and direct evaluations. Disable scheduling when you want to preserve the test definition without creating new scheduled runs. Manual runs remain available from the test page. ### Failure and recovery notifications Enable failure notifications when repeated run execution failures should alert the account owner. **Notification threshold** controls how many consecutive runs in the `failed` state are required before AIVAX sends an alert. The default is `1`. These notifications track execution errors, not behavioral test failures. A run in the `succeeded` state completed without an execution error, but its behavioral outcome can still be `loss` or `incomplete`. Those outcomes do not count toward the failure notification threshold. When **Recovery notification** is enabled, AIVAX also notifies the account after a run completes successfully following enough consecutive execution failures to reach the configured threshold. A successfully completed run resets the consecutive-failure counter, even if its behavioral outcome is `loss` or `incomplete`. Recovery therefore means execution recovered, not that the assistant passed the behavioral checks. ### Retention Succeeded and failed runs are retained for one month. Cancelled runs are retained for one day. Export any result that must remain available beyond those periods. ## Evaluation settings | Setting | Default | Accepted values | Description | | --- | ---: | --- | --- | | `validation_criteria` | `null` | String, message part, or list of message parts | Optional requirements supplied only to the judge. They do not guide the simulated user or the gateway under test. | | `resources` | `[]` | Up to 16 `{ "type", "data" }` objects | Additional context supplied to the simulated user and judge. Use `Text` for literal `data` or `RemoteResource` for content retrieved from the URL in `data`. | | `hooks` | `[]` | Up to 16 `{ "event", "url" }` objects | External callbacks for persisted runs. Supported events are `before-test`, `after-test`, `before-inference`, `after-inference`, and `context-changed`; `before-inference` and `after-inference` require an AI Gateway. | | `profile` | `medium` | `low`, `medium`, `high` | Selects the capability and price tier used by the simulated user and judge. It does not replace the model configured on the gateway under test. | | `max_turns` | `10` | `2`–`64` | Maximum number of simulated-user turns before the run ends. | | `minimum_turns` | `1` | `1`–`63`, less than `max_turns` | First turn when the simulated user may receive the option to end the conversation. | | `allow_user_exit` | `true` | Boolean | When enabled, the simulated-user prompt exposes the conversation exit token from `minimum_turns` onward. When disabled, that option is omitted from every simulated-user prompt. | | `judge_start_turn` | `1` | `1`–`63`, less than `max_turns` | First turn evaluated by the judge. The final turn is always evaluated. | | `loss_threshold` | `0.2` | `0.01`–`0.99` | Boundary used to identify a persistently unsuccessful trajectory. | | `base_threshold` | `0.9` | `0.01`–`0.99` | Score at or above which the goal is considered reached. It must be greater than `loss_threshold`, with more than `0.1` between them. | | `user_sampling.top_k` | `0.4` | `0`–`2` | Controls how many sampled communication characteristics guide the simulated user. Higher values increase variation. | | `user_sampling.max_decay` | `0.02` | `0`–`1` | Controls how sampled user characteristics may change between turns. | Reduce `max_turns` for fast, bounded regression checks. Increase it for flows that naturally require discovery or several tool calls. `minimum_turns` and `judge_start_turn` must each be less than `max_turns`; they are otherwise independent. Delay `judge_start_turn` when early clarification is expected and intermediate scores are not useful. Disable `allow_user_exit` when only the judge or the turn budget should end the test; `minimum_turns` controls only when the simulated user sees its exit option and does not delay judge decisions. Keep a wide gap between the loss and success thresholds unless the policy has been calibrated against representative conversations. Agentic Tests bills the selected gateway inference plus simulated-user and judge usage at the rates of the selected profile. See [Pricing](https://docs.aivax.net/docs/pricing.md#agentic-tests) for current rates. ## Direct API execution Each direct evaluation consumes one unit from the same account quota as persisted runs. If that quota is exceeded, the request returns HTTP 429 before the SSE stream opens. Check the HTTP status before processing events, and use bounded retries with backoff. See [Plans and limits](https://docs.aivax.net/docs/limits.md#semantic-decision-and-agentic-test-rate-limits). Use the direct generation endpoint when an application needs to run an ephemeral test and consume its events immediately. A direct execution does **not** create a persistent test case or run in the dashboard. Authenticate with a private AIVAX API key, send `Accept: text/event-stream`, and keep the key on a trusted backend. Do not expose a private key in browser code or a distributed application bundle. [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Evaluate%20Agentic%20Test) The request accepts the same core evaluation settings as a persistent test. Use `model` for the AI Gateway slug, `goal` for the desired outcome shared with the simulated user and judge, `validation_criteria` for optional judge-only requirements, `minimum_turns` and `allow_user_exit` to control when the simulated user sees its exit option, `judge_start_turn` to schedule judge evaluation, `start` for optional initial messages, `resources` for additional `Text` or `RemoteResource` context, and `external_user_id` for an identity forwarded to gateway inference. Every Server-Sent Events message contains this envelope: ```json { "timestamp": 1786329000000, "event": { "type": "chat.start", "data": { "turn_number": 1, "max_remaining_turns": 10 } } } ``` Route messages by `event.type` and concatenate streamed content chunks in order. The examples below show the `event` object inside the SSE envelope. Lifecycle markers (`start_generation`, `end_generation`, `turn_analysis_start`, `turn_analysis_end`) carry an empty object (`data: {}`); reasoning chunks share the `{ "reasoning_content": "..." }` shape. Only content-bearing events are shown in full. - **`chat.start`** — Starts a turn and reports its number and remaining turn budget. ```json { "type": "chat.start", "data": { "turn_number": 1, "max_remaining_turns": 10 } } ``` - **`chat.user_message.start_generation`** — Marks the start of simulated-user message generation (`data: {}`). - **`chat.user_message.reasoning`** — Streams one reasoning chunk exposed by the simulated-user model. Use this for debugging only. - **`chat.user_message.content`** — Streams one simulated-user message content chunk. Concatenate consecutive chunks in arrival order. ```json { "type": "chat.user_message.content", "data": { "content": "Can you explain the refund policy?" } } ``` - **`chat.user_message.end_generation`** — Marks the end of simulated-user message generation (`data: {}`). - **`chat.user_message.end_conversation`** — Reports a permitted simulated-user exit after `minimum_turns`. This event is not emitted when `allow_user_exit` is disabled. ```json { "type": "chat.user_message.end_conversation", "data": { "reason": "simulated_user_ended_conversation" } } ``` - **`chat.assistant_message.start_generation`** — Marks the start of the selected gateway response (`data: {}`). - **`chat.assistant_message.reasoning`** — Streams one reasoning chunk exposed by the gateway model (same `reasoning_content` shape). - **`chat.assistant_message.refusal`** — Reports a refusal returned by the gateway model. - **`chat.assistant_message.tool_call`** — Reports an assistant tool call, including its ID, name, and arguments. - **`chat.assistant_message.tool_result`** — Reports a tool result, including the associated call ID, name, and content. - **`chat.assistant_message.content`** — Streams one assistant response content chunk. Concatenate consecutive chunks in arrival order. ```json { "type": "chat.assistant_message.content", "data": { "content": "Refunds are available within 30 days." } } ``` - **`chat.assistant_message.end_generation`** — Marks the end of the selected gateway response (`data: {}`). - **`chat.judge.turn_analysis_start`** — Marks the start of an evaluation against the goal and any judge-only validation criteria (`data: {}`). - **`chat.judge.turn_analysis_result_ready`** — Returns the judge reasoning, normalized score, current state, trajectory measurements, and continuation decision. `score` ranges from `0.001` to `0.999`; `pass` is false only after a persistent loss is established. ```json { "type": "chat.judge.turn_analysis_result_ready", "data": { "result": { "reasoning": "The response satisfied the requested outcome and validation criteria.", "score": 0.92, "pass": true, "should_continue": false, "state": "success", "turn_delta": 0.84, "conversation_delta": 1.0, "loss_streak": 0, "required_loss_streak": 2 } } } ``` - **`chat.judge.turn_analysis_end`** — Marks the end of the current turn evaluation (`data: {}`). - **`usage_updated`** — Reports prompt, cached prompt, and completion token usage. `role` is `user`, `assistant`, or `judge` depending on the inference that produced the usage. ```json { "type": "usage_updated", "data": { "role": "judge", "usage": { "prompt_tokens": 1240, "cached_prompt_tokens": 320, "completion_tokens": 180 } } } ``` - **`unhandled_error`** — Reports an inference error, its scope, and whether the operation will be retried. `scope` is `user_inference`, `gateway_inference`, or `judge_analysis`. ```json { "type": "unhandled_error", "data": { "error": "The inference provider is temporarily unavailable.", "scope": "gateway_inference", "will_retry": true } } ``` - **`chat.validation.end`** — Reports the final outcome and closes the evaluation. The legacy event name is preserved for compatibility. `score` is included when the final outcome follows a judge evaluation, but may be absent when the turn budget is exhausted. ```json { "type": "chat.validation.end", "data": { "outcome": "success", "reason": "baseline_reached", "state": "success", "turn_number": 3, "score": 0.92, "conversation_delta": 1.0, "loss_streak": 0 } } ``` The judge state can be `active`, `at_risk`, `success`, or `loss`. A weak turn does not immediately fail a recoverable conversation: the low score and cumulative trajectory must remain at or below the configured loss threshold for the required consecutive evaluations. Final outcomes are: | Outcome | Meaning | | --- | --- | | `success` | The judge reached `base_threshold`. A permitted simulated-user exit triggers judging but does not itself pass the test; without the threshold the outcome is `incomplete`. | | `loss` | The score and cumulative trajectory remained at or below `loss_threshold` for the required consecutive judged turns. | | `incomplete` | The conversation exhausted `max_turns` without reaching success or a persistent loss. | | `interrupted` | A validation rule stopped the evaluation before it completed. | A missing or invalid key returns `401 Unauthorized`; a public API key returns `403 Forbidden`; insufficient balance returns `402 Payment Required`; and malformed fields, unavailable gateway slugs, or invalid threshold combinations return `400 Bad Request`. An inference failure can instead arrive as an SSE event after streaming begins. To investigate an inference failure, review the [AI Gateway configuration](https://docs.aivax.net/docs/inference/ai-gateway.md) used by the test. --- Source: https://docs.aivax.net/docs/inference/voice-session.html # Voice Session Voice Session is AIVAX's low-latency, stateful voice API. It keeps the authenticated AIVAX WebSocket endpoint while connecting to a realtime inference service. Events follow the OpenAI-compatible GA Realtime JSON protocol after the connection is upgraded. Use Voice Session for natural spoken conversations with streaming audio, assistant speech transcripts, server voice activity detection (VAD), interruptions, and tool calls. Use [Audio Transcriptions](https://docs.aivax.net/docs/generations/audio-transcriptions.md) for transcription-only workloads or [Speech Generation](https://docs.aivax.net/docs/generations/speech.md) when the text to synthesize is already known. ## Connect securely Open a WebSocket upgrade request to: ```text wss://inference.aivax.net/api/v1/voice-session ``` Authenticate the upgrade with a **private** account API key: ```http Authorization: Bearer ``` The `?api-key=` query parameter is also accepted when the WebSocket client cannot set headers, but headers are preferred because URLs are commonly retained in logs and monitoring systems. Public API keys cannot open Voice Sessions. > [!WARNING] > Never place a private API key in browser JavaScript, a mobile application bundle, or a browser-visible WebSocket URL. Browsers also cannot add an `Authorization` header through the native `WebSocket` constructor. Terminate the end-user connection at your backend, then let that trusted backend open and relay the authenticated AIVAX WebSocket. [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Open%20voice%20session) ## Migrate to GA Realtime events After the WebSocket upgrade, exchange OpenAI-compatible GA Realtime JSON events directly. Do not send the legacy `session_start` envelope or depend on the former AIVAX-specific STT, TTS, WAV-segment, or playback-confirmation events. The main migration changes are: | Legacy integration | GA Realtime integration | | --- | --- | | Custom session-start message | `session.update` | | Custom STT and VAD events | Native `input_audio_buffer.speech_*` events; caller transcription is currently unavailable | | WAV response segments | Base64 PCM in `response.output_audio.delta` | | Custom playback confirmation | Native cancel and conversation-item truncation events | | Custom tool-call envelopes | Native function-call items and function-call outputs | | Flat per-minute charge | Selected model's text, audio, and image token rates | Use the `type` property to route every event. Audio remains base64-encoded inside JSON text messages; do not send binary WebSocket frames. ## Configure the session Send `session.update` as the first client event. AIVAX supports the GA Realtime session fields and adds two optional selectors: - `gateway`: an AI Gateway slug available to the authenticated account; - `model`: `gpt-realtime-2.1` or `gpt-realtime-2.1-mini`. The selector supplied in the event determines the AIVAX configuration used for the session. Configure the output voice at `session.audio.output.voice` and reasoning effort at `session.reasoning.effort`. ```json { "type": "session.update", "session": { "type": "realtime", "gateway": "", "model": "gpt-realtime-2.1", "instructions": "Help the caller complete the requested task.", "audio": { "input": { "format": { "type": "audio/pcm", "rate": 24000 }, "turn_detection": { "type": "server_vad" } }, "output": { "format": { "type": "audio/pcm", "rate": 24000 }, "voice": "marin" } }, "reasoning": { "effort": "low" } } } ``` Choose the model based on the experience you need: - `gpt-realtime-2.1` for the highest-quality realtime conversations; - `gpt-realtime-2.1-mini` for lower-cost, lighter-weight realtime workloads. You can also provide standard GA Realtime instructions, tools, and turn-detection settings in the same update. Input-transcription sessions are not currently supported, so do not configure `audio.input.transcription` or depend on `conversation.item.input_audio_transcription.*` events. Wait for the server's session confirmation event before treating the configuration as active. A voice cannot be changed after the session has already produced audio; reconnect to select a different voice. ## Stream microphone audio Capture raw mono signed 16-bit little-endian PCM (`s16le`) at 24,000 Hz, base64-encode each chunk in chronological byte order, and append it to the input buffer: ```json { "type": "input_audio_buffer.append", "audio": "" } ``` Send small chunks continuously for low latency. Do not add WAV, RIFF, or other container headers to each chunk. If automatic turn detection is enabled, server VAD determines when the user starts and stops speaking and creates responses according to the session configuration. If turn detection is disabled, control the turn explicitly with the native buffer events: 1. Send one or more `input_audio_buffer.append` events. 2. Send `input_audio_buffer.commit` when the utterance is complete. 3. Send `response.create` to request the assistant response. Use `input_audio_buffer.clear` to discard buffered audio that should not become a conversation turn. ## Receive transcripts and audio Handle the native GA Realtime event stream rather than assuming one fixed sequence. In particular: - `input_audio_buffer.speech_started` and `input_audio_buffer.speech_stopped` report server-VAD boundaries; - `conversation.item.input_audio_transcription.*` events are not currently emitted because input-transcription sessions are not supported; - `response.output_audio_transcript.delta` and `response.output_audio_transcript.done` provide the assistant's spoken transcript; - `response.output_audio.delta` carries a base64 PCM audio chunk; - `response.output_audio.done` marks the end of an output-audio stream; - `response.done` reports the final response status and usage; - `error` reports a request or session error. Decode each `response.output_audio.delta` value as raw mono signed 16-bit little-endian PCM at 24,000 Hz and queue the samples for gapless playback: ```json { "type": "response.output_audio.delta", "response_id": "", "item_id": "", "output_index": 0, "content_index": 0, "delta": "" } ``` Event fields can evolve with the GA Realtime protocol. Route by `type`, retain the identifiers needed to correlate responses and items, and ignore unknown event types or fields that your client does not use. ## Handle interruptions When the caller speaks over the assistant: 1. Stop or duck local playback when the confirmed `input_audio_buffer.speech_started` event arrives. 2. Send `response.cancel` if a response is still being generated. 3. Send the native conversation-item truncation event with the assistant item identifier and the amount of audio that was actually played. Cancellation stops generation; truncation keeps the server conversation aligned with what the caller heard. Track played audio duration in the client rather than assuming every received chunk reached the speaker. Treat already-finished response races as normal and make cancel handling idempotent. ## Use tools Declare client tools through the standard GA Realtime `tools` session field. Handle `response.function_call_arguments.done` using its `response_id`, `call_id`, `name`, and JSON-encoded `arguments`. Execute each requested function, add one native `function_call_output` conversation item per `call_id`, then send a single `response.create` after the originating `response.done` and every call from that response has an output. Tools configured as AIVAX server `InternalFunctions` execute within AIVAX and require no client action. Client-defined tools continue to be returned normally and remain the client's responsibility. Your tool handler should validate arguments, preserve call identifiers, enforce timeouts, and return a useful error result when execution fails. Do not wait for a client tool result that is being executed as an AIVAX server function. ## Billing and limits For current pricing, availability, and account limits, see [Pricing](https://docs.aivax.net/docs/pricing.md) and [Plans and Limits](https://docs.aivax.net/docs/limits.md). ## Disconnect gracefully A Voice Session lasts only for the active WebSocket. It is not resumable after disconnecting. For a normal shutdown: 1. Stop microphone capture and stop sending audio. 2. Let any response or required tool-result exchange finish, or cancel it explicitly. 3. Stop and drain local playback as appropriate for the product experience. 4. Close the WebSocket with a normal close code. 5. Release audio devices, playback queues, pending calls, and session state. Handle peer closure, network loss, authentication failure, model errors, and server shutdown as terminal session outcomes. Reconnecting starts a new session: open a new authenticated WebSocket, send a new `session.update`, and rebuild any application context your experience requires. Use bounded retry with backoff for transient failures, but do not automatically retry authentication or configuration errors. --- Source: https://docs.aivax.net/docs/inference/pipelines.html # AI Pipelines AI Gateway pipelines are the processing steps AIVAX applies before and during inference. They can add context, rewrite queries, expose tools, moderate input, route models, truncate conversations, and call external workers. Most pipelines are configured in the gateway parameters. Request-level options can override some inference parameters on direct `chat/completions` calls. ## RAG RAG links [collections](https://docs.aivax.net/docs/rag/collections.md) to an AI Gateway. The gateway controls: - Collections included in retrieval. - Maximum number of retrieved documents. - Minimum score. - Reranker name. - Whether chunk references are included. - Query strategy. When a gateway has knowledge collections and the latest user message contains text, AIVAX can retrieve matching documents before the model call. For injection strategies, the retrieved context is inserted at the beginning of the last user message. If a linked collection has its own context text, that collection context is added to system instructions. Query strategies: - `Plain`: Uses the latest user message as the search query. - `Concatenate`: Joins the latest configured number of user messages line by line and searches with the combined text. - `UserRewrite`: Rewrites recent user messages into one or more search queries using a resolver model. - `FullRewrite`: Rewrites recent user and assistant messages into one or more search queries using a resolver model. - `QueryFunction`: Adds a query function to the model. The model decides when to search the linked collections, and search results are returned as tool responses. Rewrite strategies add resolver-model cost (see [Pricing](https://docs.aivax.net/docs/pricing.md)) and latency. They are useful when users ask follow-up questions such as "what about this case?" because the resolver can turn the recent conversation into a clearer search query. Defining many RAG results increases input token usage and can increase final inference cost (see [Pricing](https://docs.aivax.net/docs/pricing.md)). Start with a small result count and raise it only when the model lacks enough evidence. ## Instructions Instruction settings shape the provider-facing prompt: - **System instructions**: Added to the system instruction set. - **Remote system instruction sources**: Fetched from configured URLs and added to the system instruction set. - **User prompt template**: Replaces `{prompt}` with each user message's text before sending it to the model. - **Assistant prefill**: Adds initial assistant content before generation when the model supports prefill. Remote instruction sources are fetched as text with a 10 MB maximum response size. Their cache duration is configurable; the default is 600 seconds. Some models do not support assistant prefill, temperature, stop sequences, or reasoning effort. Integrated model validation rejects incompatible gateway settings when those limitations are known. ## Skills Skills are on-demand instruction packs available to the model. When a gateway enables skills, AIVAX loads the account skills configured on the gateway and can expose skill-related built-in functions. Read more about [skills](https://docs.aivax.net/docs/features/skills.md). ## Multimodal pre-processing Multimodal pre-processing converts selected media content into text before the main model call. The available flags are `Image`, `Audio`, `Video`, `File`, `OtherFiles`, and `All`. Use pre-processing when the main model is text-first or when you want AIVAX to normalize media into textual context. For direct multimodal models, send the original media without pre-processing so the model can inspect it directly. Media descriptions are cached by content hash for reuse. ## Parameterization The parameterization pipeline configures model request options such as: - `temperature` - `top_p` - `presence_penalty` - `frequency_penalty` - `stop` - `max_completion_tokens` - `reasoning_effort` - `verbosity` - `seed` Request-level values can override gateway values when the endpoint supports the parameter. Some integrated models reject specific parameters, and BYOK providers may have their own restrictions. ## Context truncation The context truncation pipeline uses an approximate token count. When `ContextMaximumSize` is set and the conversation exceeds the limit, the gateway follows `ContextOverflowAction`: - `Throw`: Return an error instead of calling the model. - `Truncate`: Remove older non-system messages until the conversation fits. Truncation preserves system messages and keeps at least one user message when possible. If the remaining user message still exceeds the limit, the request fails with a message-size error. On lower plans, the effective input context may be capped even when a larger context is configured; see [Plans and limits](https://docs.aivax.net/docs/limits.md#plan-limits). ## Tool message truncation `ToolContextCount` controls how many recent tool response messages keep their original content. When it is set to a value greater than zero, older tool messages remain in the conversation but their content is replaced with: ```text [tool response truncated - call this tool again] ``` This can reduce context usage in long agentic conversations. It can also hurt chains where an old tool result remains important, so use it only when the model can safely call the tool again. ## Server-side tools Server-side tools are internal functions executed by AIVAX during inference. They can come from: - Built-in tools. - Protocol functions. - Remote protocol function sources. - MCP sources. - QueryFunction RAG. - Skills. - Optional bash environment. Server-side tool events can be streamed to clients as `servertool` updates. ## Built-in tools Built-in tools can be configured on a gateway or supplied per request with `builtin_tools`. Available built-in tool flags include: - `DateTime` — current date and time through `get_date_time`; configure `dateTimeTimeZone` in built-in tool options (default: `America/Los_Angeles`, Pacific Time). - `WebSearch` - `AdvancedWebUsage` (disabled; returns an unavailable response. See [Changelogs](https://docs.aivax.net/docs/changelogs.md).) - `OpenUrl` - `Code` - `Request` - `Calendar` - `Remember` - `GenerateWebPage` - `GenerateDocument` - `XPostsSearch` - `ImageGeneration` See [Built-in tools](https://docs.aivax.net/docs/tools/builtin-tools.md). ## MCP and protocol functions Tools listed by an MCP source become available to the model with their declared schemas. MCP tool results can include text, images, and audio; media results are attached to the conversation as additional messages when supported. Protocol functions expose HTTP callbacks or AIVAX callback URLs as model-callable tools. Remote protocol function sources are fetched and cached before their tools become available to the model. See [Protocol functions](https://docs.aivax.net/docs/tools/protocol-functions.md) and [MCP](https://docs.aivax.net/docs/tools/mcp.md). ## Function interpreter A tool handler can add tool-calling behavior for models that do not reliably produce native tool calls. Supported values are: - `native` or `null`: Use native model tool calling. - `react.v1.selfcall`: Use the ReAct-style self-call handler. If a handler name is not recognized, gateway configuration fails at inference time. ## Moderation Moderation is an input gate that runs before the main model. When at least one moderation category is enabled, AIVAX sends the available textual conversation to a safeguard model. The safeguard evaluates the latest user request in the context established by the conversation and returns a score from 0 to 10 for every category. | Category | Gateway property | What the safeguard evaluates | | --- | --- | --- | | **Violence and hate speech** | `violenceThreshold` | Violence, hate, extremism, threats, or encouragement of physical harm. | | **Sexual and explicit content** | `sexualExplicitThreshold` | Sexually explicit or adult content. | | **Political topics** | `politicalThreshold` | Political persuasion, campaigning, manipulation, or highly political content. | | **Dangerous content** | `dangerousContentThreshold` | Weapons, explosives, cyber abuse, self-harm, or other dangerous acts and instructions. | | **Jailbreak attempts** | `jailbreakThreshold` | Attempts to override instructions, reveal protected instructions, exfiltrate data, or inject prompts. | ### Understand sensitivity levels The value configured in the gateway is a **sensitivity level**, not the safeguard score itself. A higher level lowers the score required to block the input. | Sensitivity level | Safeguard scores that block | | ---: | --- | | `0` | Category disabled | | `1` | `10` | | `3` | `8`–`10` | | `5` | `6`–`10` | | `8` | `3`–`10` | | `10` | `1`–`10` | For enabled categories, the blocking cutoff is `11 - sensitivity level`. A safeguard score of `0` never blocks. Configure each category independently; the request is blocked when any enabled category reaches its cutoff. ### Add gateway-specific rules **Additional moderation rules** let a gateway administrator describe policy that is not fully expressed by the built-in category descriptions. The safeguard reads these rules together with the built-in policy and uses them to calibrate all five scores. For example: ```text Allow users to describe accidents and injuries when asking for insurance coverage or assistance. Treat requests for instructions to cause an accident or harm someone as dangerous content. ``` Additional rules guide classification; they do not create a separate score or block an input directly. At least one category must have a sensitivity level above `0` for moderation to run. In the example above, enable **Dangerous content** so the safeguard score can produce a blocking decision. Write rules as short policy statements with explicit allowed and disallowed cases. Do not include secrets, credentials, or private operational data because the rules are stored with the gateway configuration and sent to the safeguard model during moderation. The following gateway fragment enables different sensitivity levels and gives the safeguard domain-specific guidance. The values are an example, not a recommended production baseline: ```json { "moderationParameters": { "violenceThreshold": 4, "sexualExplicitThreshold": 4, "politicalThreshold": 2, "dangerousContentThreshold": 6, "jailbreakThreshold": 7, "additionalRules": "Allow descriptions of accidents and injuries for insurance support. Treat instructions to cause accidents or harm someone as dangerous content." } } ``` ### What happens when an input is blocked When any enabled category reaches its cutoff: 1. AIVAX marks the original conversation messages as unavailable to the main inference request. 2. AIVAX replaces them with an instruction identifying the categories that caused the block. 3. The main model generates a refusal instead of answering the original request. The refusal is model-generated; moderation does not return a fixed response body. If the safeguard cannot produce a valid moderation result, the request fails before the normal completion is generated. ### Context and current limitations The safeguard receives the available conversation history, not only the latest message. Message roles and text are preserved as untrusted serialized conversation data so instructions inside the conversation cannot replace the safeguard policy. If the conversation exceeds the safeguard context window, older context can be truncated. Moderation currently applies only to input text: - Generated output is not moderated. - Images, audio, video, and file contents are not analyzed. The safeguard receives only a marker indicating that media was present. - Tool or worker authorization still requires application-level policy; moderation is not an authorization mechanism. - Moderation adds a safeguard inference before the main inference, which adds latency and billable moderation usage. Use moderation for broad safety policy. Use workers when the decision depends on external identity, account state, or business-specific policy. ## Workers Configure worker events and endpoint details in the gateway parameters; implement the event behavior at your external endpoint. See [AI Workers](https://docs.aivax.net/docs/inference/workers.md). --- Source: https://docs.aivax.net/docs/inference/structured-responses.html # Structured Responses AIVAX can produce structured JSON through two paths: - `response_schema`: AIVAX validates the final model output against a JSON Schema and retries with validation feedback when the output is invalid. This is the JSON Healing path. - `response_format`: AIVAX passes a native OpenAI-compatible response format to the provider, or applies healing when `healing_options` is present or automatic JSON Healing is enabled on the account. Use structured responses when another system will consume the output and free-form text would be fragile. ## How JSON Healing works When `response_schema` is present, AIVAX adds schema instructions to the model request. After the model generates an answer, AIVAX tries to extract JSON from: - The complete generated text. - Common repaired variants, such as missing opening or closing braces. - JSON code blocks found in the generated text. If the extracted JSON does not validate against the schema, AIVAX adds a feedback message with the validation errors and asks the model to generate the JSON again. This continues until a valid JSON value is generated or the configured attempt limit is reached. JSON Healing improves reliability, but it is not an absolute guarantee. If the model repeatedly fails the schema, the request can fail after the attempt limit. ## Choosing a mode Use `response_schema` when AIVAX should own validation and repair. This is the safest mode for models that do not natively support structured output, for responses that use tools before producing JSON, and for systems that cannot tolerate malformed JSON. Use `response_format` with `type: "json_schema"` when the provider model should handle the structured output natively. AIVAX will still use the schema internally, and healing is applied when `response_format.json_schema.healing_options` is provided or the account has automatic JSON Healing enabled. Use `json_only: true` when the HTTP response body should be only the final JSON. This removes the normal chat completion envelope, choices, usage, and generation metadata from the response body. For a troubleshooting checklist covering refusals, incomplete streams, and native enforcement versus healing, read [How to fix invalid JSON from an LLM API with structured outputs](https://aivax.net/blog/structured-output-healing-boundary/). ## Basic example Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Inference%20(chat%20completions))
POST /v1/chat/completions
```json { "model": "@google/gemini-2.5-flash", "prompt": "Search for recent news about electric vehicles.", "stream": true, "builtin_tools": { "tools": [ "WebSearch" ], "options": { "web_search_mode": "full" } }, "response_schema": { "type": "object", "properties": { "news": { "type": "array", "items": { "type": "object", "properties": { "title": { "type": "string", "description": "News title" }, "summary": { "type": "string", "description": "News summary" } }, "required": ["title", "summary"] } } }, "required": ["news"] } } ``` `builtin_tools` is optional. If you enable tools, the selected model must support tool calling or the gateway must provide a tool handler. ## Native structured output Use `response_format` when you want to use the provider's native JSON Schema support: ```json { "model": "@openai/gpt-4o", "messages": [ { "role": "user", "content": "List 3 European capitals." } ], "response_format": { "type": "json_schema", "json_schema": { "schema": { "type": "object", "properties": { "capitals": { "type": "array", "items": { "type": "object", "properties": { "city": { "type": "string" }, "country": { "type": "string" } }, "required": ["city", "country"] } } }, "required": ["capitals"] } } } } ``` If the selected integrated model has strict JSON support, AIVAX can forward the schema using the provider's JSON Schema response format. ## Enabling healing in `response_format` You can explicitly enable JSON Healing inside `response_format.json_schema`: ```json { "model": "@openai/gpt-4o", "messages": [ { "role": "user", "content": "Return a short status object." } ], "response_format": { "type": "json_schema", "json_schema": { "schema": { "type": "object", "properties": { "status": { "type": "string" }, "message": { "type": "string" } }, "required": ["status", "message"] }, "healing_options": { "max_attempts": 5 } } } } ``` `max_attempts` must be between 1 and 10. Each retry is another model generation attempt and can increase cost and latency. If healing frequently reaches the attempt limit, adjust the instruction and schema before raising the limit. Common causes are overly rigid schemas, missing `required` fields, vague prompts, noisy tool results, or a model that is too small for the task. ## `json_only` mode Set `json_only: true` when the client should receive only the generated JSON: ```json { "model": "@openai/gpt-4o", "messages": [ { "role": "user", "content": "List 3 European capitals." } ], "response_schema": { "type": "object", "properties": { "capitals": { "type": "array", "items": { "type": "object", "properties": { "city": { "type": "string" }, "country": { "type": "string" } }, "required": ["city", "country"] } } }, "required": ["capitals"] }, "json_only": true } ``` With `stream: false`, the HTTP response body is the final JSON with `Content-Type: application/json`. With `stream: true`, AIVAX sends the complete JSON as one SSE data event and then sends `[DONE]`. ## Supported schema features AIVAX validates the generated JSON with JSON Schema. The documented supported features are: - `string`: `minLength`, `maxLength`, `pattern`, `format`, and `enum`. - `number` and `integer`: `minimum`, `maximum`, `exclusiveMinimum`, `exclusiveMaximum`, and `multipleOf`. - `array`: `items`, `uniqueItems`, `minItems`, and `maxItems`. - `object`: `properties` and `required`. - `boolean` and `bool`. - `null`. - Multiple types, for example `"type": ["string", "number"]`. Use `required` for fields your application must receive. Use explicit `items` for arrays. Use `enum`, `format`, and length or pattern restrictions when the accepted values are known. ## Practical patterns For extraction, tell the model which fields must be inferred, which fields should be `null` when missing, and which fields must not be invented. For classification, use `enum` for the final label and add a short `reason` field when humans need to audit the decision. For tool-backed JSON, allow the model to use tools before producing the JSON and include source fields when the application needs provenance. For database writes or external API calls, validate the JSON again in your application. AIVAX validates the JSON shape, but your application remains responsible for business rules such as permissions, valid IDs, date ranges, and account-specific constraints. --- Source: https://docs.aivax.net/docs/inference/workers.html # AI Workers AI Gateway workers are HTTP hooks that let an external service control gateway execution at runtime. A worker can allow an event, stop it, rewrite context, add instructions or tools, or replace a server-side tool result. Use workers when a rule must be decided outside the prompt. Common cases include subscription checks, CRM enrichment, tenant-specific policy, audit logging, dynamic tool blocking, and replacing a model-visible tool result with data from an internal system. Workers run in the inference critical path. Every worker event adds an HTTP request before the gateway can continue, so the endpoint must respond quickly and predictably. ## Request format When a configured worker event fires, AIVAX sends a `POST` request to the gateway's worker URL. ```json { "gatewayId": "your-gateway-id", "moment": "2025-12-29T17:04:39", "event": { "name": "message.received", "data": { "messages": [ { "role": "system", "content": "User local date is Monday, December 29, 2025 (timezone is America/Sao_Paulo)" }, { "role": "user", "content": "Good morning" } ], "origin": "ChatCompletionsApi", "externalUserId": "customer-123", "metadata": {} } } } ``` The exact `event.data` shape depends on the event. Always validate `gatewayId` when one endpoint serves more than one gateway. ## Authentication When the account has a hook key, AIVAX sends `X-Request-Nonce`. The nonce is a BCrypt hash derived from the account hook key. Validate this header before trusting the body, especially when the worker releases private data, changes context, or authorizes tool usage. Treat `externalUserId`, `metadata`, messages, and tool arguments as untrusted input. ## Response behavior After sending the request, AIVAX handles the worker response as follows: | Response | Behavior | |---|---| | `Content-Type: application/json+worker-action` | Execute the action described in the JSON body. | | `2xx` without `application/json+worker-action` | Continue normally. | | Non-OK response without `application/json+worker-action` | Stop the event. | If the worker request fails with an HTTP request exception, AIVAX logs the failure and stops the event. Choose fail-open or fail-closed behavior intentionally. Return `2xx` when enrichment is optional. Return a non-OK response when authorization, compliance, or business policy cannot fail open. ## `message.received` The `message.received` event fires after the gateway prepares the incoming message context and before the model call. ```json { "name": "message.received", "data": { "messages": [], "origin": "ChatCompletionsApi", "externalUserId": "customer-123", "metadata": {} } } ``` To modify the context, return `Content-Type: application/json+worker-action` with `type: "message.received.response"`: ```json { "type": "message.received.response", "data": { "rewrites": [ { "type": "add-system", "message": "Answer in formal English." } ] } } ``` Available rewrite actions: | Action | Description | Parameters | |---|---|---| | `clear` | Removes context elements. | `argument`: `messages`, `meta`, `system`, `tools`, `skills`, `all`, or omitted. | | `add-message` | Adds a message to the conversation. | `message`: OpenAI-compatible message object. | | `remove-message` | Removes a message by index. | `index`: zero-based message index. | | `add-system` | Adds a system instruction. | `message`: instruction text. | | `add-tool` | Adds an OpenAI-compatible tool definition. | `tool`: tool JSON object. | | `add-protocol-tool` | Adds a [protocol function](https://docs.aivax.net/docs/tools/protocol-functions.md). | `tool`: protocol function definition. | | `add-mcp-source` | Adds the tools discovered from an [MCP](https://docs.aivax.net/docs/tools/mcp.md) source to the context. | `source`: MCP source object with `url`, `headers`, `name`, and/or `cacheDuration`. | ### Replace the user context ```json { "type": "message.received.response", "data": { "rewrites": [ { "type": "clear" }, { "type": "add-message", "message": { "role": "user", "content": "The original message was removed by an external policy check. Tell the user they need an active subscription to continue." } } ] } } ``` ### Remove a message ```json { "type": "message.received.response", "data": { "rewrites": [ { "type": "remove-message", "index": 0 } ] } } ``` ### Add a temporary MCP source ```json { "type": "message.received.response", "data": { "rewrites": [ { "type": "add-mcp-source", "source": { "name": "Internal CRM", "url": "https://crm.example.com/mcp", "headers": { "Authorization": "Bearer " }, "cacheDuration": 600 } } ] } } ``` Use `add-mcp-source` when the tool list needs to depend on the message, user, channel, or an external policy. AIVAX lists the tools from the MCP server, converts each schema into a model-callable function, and makes those tools available only for that inference. For permanent sources, configure MCP directly in the AI Gateway. ## `tool.called` The `tool.called` event fires before AIVAX executes an internal server-side tool. ```json { "name": "tool.called", "data": { "toolName": "check_order", "toolArguments": { "order_id": "A123" }, "origin": "ChatCompletionsApi", "externalUserId": "customer-123", "metadata": {} } } ``` Return a non-OK response to block the tool call. Return `2xx` to let AIVAX execute the tool normally. To replace the tool result, return `Content-Type: application/json+worker-action` with `type: "tool.called.response"`: ```json { "type": "tool.called.response", "data": { "result": "Order A123 is paid and scheduled for delivery tomorrow.", "messages": [] } } ``` Fields of `data`: | Field | Description | |---|---| | `result` | Textual content injected as the tool result. | | `messages` | Optional additional OpenAI-format messages attached to the conversation context. | When `tool.called.response` is returned, AIVAX uses the worker-provided result instead of executing the default tool handler. ## Example: blocking unauthorized users The example below shows a Cloudflare Worker that blocks a gateway call when the external user is not allowed. ```js export default { async fetch(request, env) { if (request.method !== "POST") { return new Response("Method not allowed", { status: 405 }); } const body = await request.json(); if (body.gatewayId !== env.CHECKING_GATEWAY_ID) { return new Response(); } if (body.event?.name !== "message.received") { return new Response(); } const externalUserId = body.event.data.externalUserId; const allowedUsers = new Set((env.ALLOWED_USERS || "").split(",")); if (!allowedUsers.has(externalUserId)) { return new Response("User is not authorized", { status: 403 }); } return new Response(); } }; ``` ## Example: replacing a tool with an internal system Use `tool.called` when the model should see a tool result, but the actual data should come from your system. ```js export default { async fetch(request, env) { const body = await request.json(); if (body.event?.name !== "tool.called") { return new Response(); } const { toolName, toolArguments, externalUserId } = body.event.data; if (toolName !== "check_order") { return new Response(); } const orderId = toolArguments?.order_id; const orderResponse = await fetch(`${env.INTERNAL_API}/orders/${orderId}`, { headers: { "Authorization": `Bearer ${env.INTERNAL_API_TOKEN}` } }); if (!orderResponse.ok) { return new Response(JSON.stringify({ type: "tool.called.response", data: { result: `The order ${orderId} could not be retrieved for user ${externalUserId}. Ask the user to confirm the order number.` } }), { headers: { "Content-Type": "application/json+worker-action" } }); } const order = await orderResponse.json(); return new Response(JSON.stringify({ type: "tool.called.response", data: { result: `Order ${order.id}: status ${order.status}, estimated delivery ${order.eta}.` } }), { headers: { "Content-Type": "application/json+worker-action" } }); } }; ``` This pattern prevents exposing the internal API directly to the model. The worker remains responsible for authenticating the request, validating the user, calling the internal system, and deciding how much data can return to the model context. For how workers fit into gateway execution, see [Pipelines](https://docs.aivax.net/docs/inference/pipelines.md). --- Source: https://docs.aivax.net/docs/web-foundation/web-search.html # Web Search Web Search retrieves current information from the internet for research, fact checking, and answers that need sources beyond the model's training data. Use it to discover relevant pages; use [Fetch and OCR](https://docs.aivax.net/docs/web-foundation/fetch-and-ocr.md) when you already have a URL or need to read a source in more detail. ## Choose how to use Web Search | Integration | When to use it | | --- | --- | | [Built-in tools](https://docs.aivax.net/docs/tools/builtin-tools.md) | Let an AIVAX model decide when to search during inference. Enable `WebSearch` in the gateway or request's `builtin_tools` configuration. | | [Web utilities MCP](https://docs.aivax.net/docs/mcp-utilities/web-utilities-mcp.md) | Give an MCP-compatible agent, IDE, or automation client access to the `web_search` tool without running an AIVAX model inference. | | Direct API | Call the search endpoint from your backend when no model or agent is involved — scheduled monitors, dataset enrichment, or pre-fetching context before inference. | `AdvancedWebUsage` is disabled and returns an unavailable response. See [Changelogs](https://docs.aivax.net/docs/changelogs.md) for details. ## Search and verify sources Write a focused query that includes the topic and any relevant date, product version, or location. For several independent questions, use separate searches rather than combining unrelated topics into one query. For direct API requests, narrow results with `country` (a two-letter country code), `language` (a language code), and `includeDomains` (trusted domains). Set `topn` to request up to 25 results when recall matters, such as when surveying competing sources; otherwise, keep the default small count. For parameters of the built-in Web Search tool, see its [reference](https://docs.aivax.net/docs/tools/builtin-tools.md). Search results help locate evidence; they do not guarantee that a source is accurate or current. Check publication dates, prefer primary sources, and fetch the relevant pages before relying on details that a result summary may omit. Keep source links with the answer so readers can verify the claims. Treat retrieved text as external, untrusted content, not as instructions for your agent. ## API reference For the supported request, response, authentication, and error contract, use the API Reference: [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Search%20the%20web) ## Pricing and limits Web Search is billed per search. Calls through [Web utilities MCP](https://docs.aivax.net/docs/mcp-utilities/web-utilities-mcp.md) use the same pricing as the corresponding built-in tool and are charged to the authenticated account. Model inference, when used, is billed separately. See [Pricing](https://docs.aivax.net/docs/pricing.md) for current Web Search and advanced search charges, and [Plans and limits](https://docs.aivax.net/docs/limits.md) for account quotas and rate limits. --- Source: https://docs.aivax.net/docs/web-foundation/fetch-and-ocr.html # Fetch and OCR Fetch and OCR extracts readable text from web pages and supported documents so applications and agents can use their content for summaries, analysis, or knowledge workflows. The Fetch API can also convert the extracted text into structured JSON using a schema you provide. Use [Web Search](https://docs.aivax.net/docs/web-foundation/web-search.md) first if you need to discover sources rather than read a known URL. Web pages are processed to remove markup and non-content elements. Document extraction and optical character recognition (OCR) make supported non-text content available as text. Review extracted content before relying on it: scan quality, complex layouts, and tables can affect the result. ## What you can extract | Content | Supported formats | Extracted content | | --- | --- | --- | | Web pages | HTML, XHTML | Readable page content, such as articles, documentation, and product information, with markup and non-content elements removed. JavaScript and CSS is rendered before extraction. | | Plain text and Markdown | TXT, Markdown | The document's text, including existing Markdown formatting. | | PDFs | PDF | Text from digital documents and OCR text from scanned pages. Mixed PDFs can combine direct text extraction with OCR where needed. | | Images containing text | PNG, JPEG, WebP, TIFF, BMP | Recognized text from screenshots, scanned documents, receipts, and other images with legible writing. | | Word-processing documents | DOC, DOCX, ODT, RTF | Document text converted into a readable textual representation. | | Presentations | PPT, PPTX, ODP | Textual slide content. | | Spreadsheets and tabular files | XLSX, ODS, CSV | Cell and row content as text, for downstream reading or analysis. | | E-books | EPUB | The publication's text. | For example, you can fetch an online article, read a PDF manual, extract text from a receipt image, or turn a spreadsheet into text for an agent to analyze. Provide a URL that returns the actual page or file, not a sharing page that requires a login. When submitting a data URI through the API, declare the content's correct MIME type. The result is extracted text, not a pixel-perfect copy of the original document. Do not assume that layout, table structure, charts, or embedded images will be reproduced exactly. Spreadsheet extraction does not execute formulas or macros. ## Fetch API vs. Media Descriptions Use the **Fetch API to extract existing text**. Use [Media Descriptions](https://docs.aivax.net/docs/generations/media-descriptions.md) to **interpret media with AI**, optionally guided by what your application needs to learn from it. | | Fetch API | Media Descriptions | | --- | --- | --- | | Main purpose | Retrieve readable text from pages and documents, using OCR for supported images and scanned PDFs. | Generate descriptions or extract information from images, PDFs, audio, and video using AI. | | Output | Extracted text, optional schema-guided JSON generated from that text, processing-unit usage, and per-item errors. | Model-generated content focused by your guidance, which can describe visual or audiovisual information beyond the text present in the source. | | Images and PDFs | Read text, such as the words on a receipt or the paragraphs in a manual. | Describe visual content or interpret a document, such as explaining a diagram or identifying information relevant to a question. | | Audio and video | Not an audio/video understanding or transcription API. | Analyze audio and video content. For a dedicated speech-to-text workflow, use [Audio Transcriptions](https://docs.aivax.net/docs/generations/audio-transcriptions.md). | | Billing | Extraction processing units (PUs), with plan-dependent allowances and rates. Optional JSON conversion is billed separately and is not covered by the extraction allowance. | AI usage charges under Media Descriptions pricing; Fetch PU allowances do not replace these charges. | For a PDF report, choose Fetch when you need its text for indexing or later analysis. Choose Media Descriptions when you need an explanation of its charts or a guided interpretation of the content. For a receipt image, Fetch reads the printed text; Media Descriptions can interpret the receipt according to your extraction guidance. Neither guarantees perfect results. Fetch can lose text or structure because of OCR and layout limitations. Media Descriptions can omit details or introduce incorrect interpretations because its output is model-generated. Verify consequential details against the original source and compare the billing paths in the pricing section below. ## Choose an integration | Integration | When to use it | | --- | --- | | Fetch API | Your application controls which URLs or inline base64 data URIs to process and needs structured results, processing-unit usage, and per-item errors. | | [Web utilities MCP](https://docs.aivax.net/docs/mcp-utilities/web-utilities-mcp.md) | An MCP-compatible client needs to read public URLs through `fetch_url`. This tool accepts one to five URLs per call. | | [Built-in tools](https://docs.aivax.net/docs/tools/builtin-tools.md) | An AIVAX model needs to read a URL during inference. Enable `OpenUrl` and follow the URL Context configuration guide. | The integrations have different input contracts. In particular, the MCP tool accepts public URLs; use the Fetch API for inline base64 data URIs. ## Fetch content with the API Authenticate with an AIVAX API key; see [Authentication](https://docs.aivax.net/docs/authentication.md). Supply a non-empty `contents` array containing URLs or base64 data URIs. Each item is limited to 10 MB. The `system.v1.web.fetch` operation accepts these JSON request fields: | Field | Type | Required | Behavior | | --- | --- | --- | --- | | `contents` | Array of strings | Yes | Non-empty list of URLs or base64 data URIs to extract. | | `returnErrors` | Boolean | No | Defaults to `true`: includes a result for each failed item. When `false`, failed items are omitted. | | `responseSchema` | JSON Schema object | No | Converts each item's extracted text into JSON guided by this schema. Omit it or use `null` for text-only extraction. The same schema applies to every item in the batch. | | `responseSchema.instructions` | String | No | Optional extraction guidance inside the schema, such as which details to select or how to handle missing information. This is an AIVAX extension, not a standard JSON Schema keyword. | ### Extract structured JSON Provide `responseSchema` when your application needs fields from the source rather than only its text. Describe the expected properties, types, and required fields with JSON Schema. Use property descriptions and optional `responseSchema.instructions` to clarify what to extract. For example, an object schema can request a receipt's merchant, date, and total; the embedded reference includes a complete JSON request example. JSON conversion runs after text extraction. Successful results retain `extractedText` alongside `extractedObject`, so you can compare the generated fields with the extracted source. Without a schema, no JSON conversion runs. If extraction or JSON conversion fails, the item follows `returnErrors`; a JSON conversion failure does not return a text-only success. ### Read the results The response contains a `results` array with these fields: | Field | Meaning | | --- | --- | | `index` | Zero-based position in the input `contents` array. Use it to match results to inputs, especially when failed items are omitted. | | `extractedText` | Readable source text, retained on successful JSON conversion; `null` for a failed item. | | `extractedObject` | Generated JSON value when `responseSchema` is supplied; otherwise `null`. Also `null` for a failed item. | | `processingUnits` | PUs for fetching and text/OCR extraction, separate from JSON conversion. | | `jsonProcessingUnits` | PUs for JSON conversion; `0` when no schema is supplied. | | `error` | Error message for a failed item when `returnErrors` is `true`; `null` on success. Failed items report both PU fields as `0`. | [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Fetch%20web%20contents) The embedded reference defines the current request and response contract. Check individual results before passing their text to the next step; a failed extraction is not evidence that the source contains no relevant information. ## Access and extraction limitations A destination may block automated access or require authentication. Fetching a URL does not bypass access restrictions. Use accessible sources or content you are authorized to submit, and correct invalid or inaccessible inputs before retrying. Treat extracted text as untrusted source material, not as instructions for your application or agent. OCR output may need manual review for exact numbers, names, or other consequential details. Schema-guided JSON is model-generated from that text, not from a fresh visual interpretation of the source. It can inherit extraction errors or contain incorrect values; validate its structure and verify consequential fields against the original. ## Pricing and limits Fetch and OCR extraction is metered in `processingUnits`. Daily included allowances and rates for uncovered extraction depend on the account plan. Each extraction is either fully covered or billed in full; coverage is not split within one extraction. Optional JSON conversion is metered separately in `jsonProcessingUnits`, with a variable inference-based PU price and no coverage from the daily extraction allowance. Do not add both counts and apply the OCR rate to the total. See [Pricing](https://docs.aivax.net/docs/pricing.md#web-search-ocr-and-fetch) for allowances and charges rather than estimating cost from extracted text length. [Web utilities MCP](https://docs.aivax.net/docs/mcp-utilities/web-utilities-mcp.md) uses the same pricing as the corresponding built-in tool. Model inference, when used to analyze the extracted content, is billed separately. See [Plans and limits](https://docs.aivax.net/docs/limits.md) for account quotas and rate limits. The Fetch API requires a positive account balance. --- Source: https://docs.aivax.net/docs/generations/decisions.html # Semantic decisions Semantic decisions evaluate named questions against a shared state and return structured answers rather than a generated explanation. Use them to route support requests, select a category, check a condition, or assign an ordered score. A request supplies the model, the evidence in `state`, and a `questions` object. Each question has an ID you choose; the response uses the same ID in `answers`. You can ask different question types in one request without making separate API calls. Use [structured responses](https://docs.aivax.net/docs/inference/structured-responses.md) instead when you need a larger generated object or a written explanation. For embedding-based similarity between documents and labels, see [Text classification](https://docs.aivax.net/docs/rag/classification.md). ## Choose a model All models below support `choice`, `noul`, and `score`. See [Pricing](https://docs.aivax.net/docs/pricing.md#semantic-decisions) for model rates and [Plans and limits](https://docs.aivax.net/docs/limits.md#semantic-decision-model-limits) for model-specific limits. | Model | | --- | | `@supersonic-labs/julia-1` | | `@typesafe/jev-1.13` | | `@respan/span-01` | | `@respan/span-01-lite` | | `@jaredpalmer/kev-4b` | `@typesafe/jev` is also accepted and currently resolves to `@typesafe/jev-1.13`. An unspecified context limit does not mean unlimited input. Model-specific limits and interpretation of scores can differ; validate a model on representative examples before switching production traffic. An authenticated API key and a positive account balance are required, including when selecting a model with a zero base token price. See [Authentication](https://docs.aivax.net/docs/authentication.md), [Pricing](https://docs.aivax.net/docs/pricing.md), and [Plans and limits](https://docs.aivax.net/docs/limits.md). ### Discover models programmatically `GET /api/v1/information/decisions-models.json` lists the current decision-model catalog without authentication. The response's `data` array contains: - `name`: the canonical identifier to use in a decision request. - `aliases`: other accepted identifiers for that model. - `contextLength`: the advertised context in tokens, or `null` when unspecified. - `releaseDate`: the catalog release date in `yyyy-MM-dd` format. - `capabilities`: supported question types (`noul`, `choice`, and/or `score`). - `inputPricePerMillionTokens` and `outputPricePerMillionTokens`: base USD prices, before account and plan adjustments. Use this listing to populate model selectors rather than maintaining a separate hardcoded catalog. It describes configured models, not a real-time provider health check. Model-specific option and question limits are not included in this listing. ## Write the questions | Type | `criteria` format | Use it for | | --- | --- | --- | | `choice` | Object mapping choice IDs to descriptions | Select one destination or category | | `noul` | Object with nonempty `false` and `true` descriptions | Evaluate a Boolean condition | | `score` | Ordered array of level descriptions | Evaluate a graded condition | Every question requires nonempty `instructions`. Use descriptions that distinguish the options, not just opaque IDs. For example, `"billing": "Charges, payments, and refunds"` provides more evidence than `"billing": "B"`. For `noul`, both criteria are required. Describe what counts as false as carefully as what counts as true, especially when the state might omit the relevant information. For `score`, keep the level order consistent across requests. ## Evaluate a support request Send the following JSON body to `POST /api/v1/generations/decisions`. The example uses Julia-1; its state may be text, an object, or an array. ```json { "model": "@supersonic-labs/julia-1", "state": { "message": "I was charged twice. Please return the extra payment." }, "questions": { "department": { "type": "choice", "instructions": "Which team should handle this request?", "criteria": { "billing": "Charges, payments, and refunds", "technical": "Software errors and service outages", "sales": "Plans, pricing, and purchases" } }, "refund_requested": { "type": "noul", "instructions": "Does the customer explicitly ask for money back?", "criteria": { "false": "The customer does not ask for money to be returned", "true": "The customer asks for a refund or return of a payment" } }, "urgency": { "type": "score", "instructions": "How urgent is the request based on the stated deadline?", "criteria": [ "No deadline stated", "A deadline is stated, but it is not today", "The customer explicitly needs resolution today" ] } } } ``` Keep only relevant evidence in the state. Instructions should explain the decision, not ask for a chain of reasoning or additional text. Avoid overlapping choice descriptions unless that ambiguity is intentional. ### Read the response The success response is a JSON object directly, **without a `data` envelope**. It contains: - `id`: the decision identifier. - `model`: the canonical model identifier used for the request. - `provider`: the provider reported for the result. - `answers`: an object keyed by your question IDs. - `usage`: `input_tokens`, `output_tokens`, and the billed `cost`. For Julia-1, an illustrative `answers` object for the example above is shown below. These numbers explain the format; they are not a recorded response or a quality guarantee. ```json { "department": { "type": "choice", "choice": "billing", "probabilities": { "billing": 0.90, "technical": 0.06, "sales": 0.04 } }, "refund_requested": { "type": "noul", "noul": 0.95, "probabilities": { "false": 0.05, "true": 0.95 } }, "urgency": { "type": "score", "score": 0.3, "legend": { "0": "No deadline stated", "1": "A deadline is stated, but it is not today", "2": "The customer explicitly needs resolution today" }, "probabilities": { "0": 0.8, "1": 0.1, "2": 0.1 } } } ``` With Julia-1: - `choice` is the selected caller-defined ID, not the option description. - `noul` is the probability assigned to the true criterion, not a JSON Boolean. Your application chooses the threshold and how to handle uncertain cases. - `score` is the expected **zero-based level index**. In the example, `0 × 0.8 + 1 × 0.1 + 2 × 0.1 = 0.3`. It is not necessarily an integer, and it is not a normalized 0–1 score when there are more than two levels. - `probabilities` are keyed by choice IDs, `false`/`true`, or score-level indices. `legend` describes the score levels. Other models may omit optional fields such as `probabilities`, `legend`, or `confidence`. Do not assume that every provider uses the same score scale or confidence definition. A high probability is not proof that the decision is correct; validate thresholds and escalation rules with labeled examples from your own domain. ## Account rate limits Semantic decision requests share an account-level rate limit across models and API keys. Each request counts once, even when it contains multiple questions. This quota is separate from the daily subscription allowance and applies to both included and paid usage. See [Plans and Limits](https://docs.aivax.net/docs/limits.md#plan-limits) for the Free, Pro, and Max thresholds. Requests above the limit return `429 Too Many Requests` before evaluation. Pace calls across the account and retry with backoff after the rate-limit window clears; changing API keys within the same account does not provide a separate quota. ## Model-specific limits See [Semantic decision model limits](https://docs.aivax.net/docs/limits.md#semantic-decision-model-limits) for the current context, question, option, and payload limits. Limits interact: shorten descriptions or reduce the option count instead of assuming every maximum can be used at once. Inputs exceeding the context or question/options budget are rejected, not silently truncated. The literal `` is reserved and cannot appear in Julia-1 state, instructions, or option descriptions. ## Usage and cost For Julia-1, input usage sums the encoded sequence for each question, excluding padding. The shared state is therefore counted again for each question. Multiple questions over one state do not have the same input usage as one question over that state. Julia-1 does not generate text, so its `output_tokens` value is zero. Julia-1 is currently eligible for the daily semantic decision allowance on Free, Pro, and Max. Other decision models are billed normally. The allowance is shared across eligible decision calls, not reserved for each question or API key. See [Plans and Limits](https://docs.aivax.net/docs/limits.md#included-daily-subscription-allowances) for relative plan capacity and coverage rules. When not covered, Julia-1 input is billed at the [listed rate](https://docs.aivax.net/docs/pricing.md#semantic-decisions), subject to account and plan adjustments. Use the returned `usage.cost` for the actual billed amount; it is zero when the input is fully covered by the allowance. ## Errors and reliable use - **Invalid model or question:** check the exact model identifier, question type, instructions, and criteria shape. Question IDs and choice IDs must be nonempty. - **Context or option limit exceeded:** shorten the state or descriptions, reduce the option count, or select a model with suitable limits. Retrying the same invalid input will not resolve it. - **Authentication or balance error:** verify the API key and account balance before retrying. A zero-priced model still requires a positive balance. - **Rate limit (429):** reduce the account's request rate and retry with backoff. Multiple questions in one request still count as one request, but model-specific payload limits and per-question usage remain applicable. - **Temporary capacity or provider unavailability:** avoid an immediate parallel retry storm. Reduce concurrency and use bounded retries with backoff for transient failures. Evaluate `choice`, `noul`, and `score` separately when validating a model: success at routing does not establish reliable scoring or Boolean behavior. Include ambiguous and incomplete states in your test set, and use human review where a wrong decision has material consequences. A retry is a new request; do not assume automatic deduplication or identical model outputs. ## API reference [API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Evaluate%20semantic%20decisions) --- Source: https://docs.aivax.net/docs/generations/speech.html # Speech Generation Use Speech Generation when your application already has final text and needs playable audio without running a chat completion. Typical uses include narrating an article or notification, voicing an IVR prompt, producing a draft voice-over for review, or generating audio files for offline playback. Authenticate requests with an AIVAX API key. See [Authentication](https://docs.aivax.net/docs/authentication.md) for authorization guidance. ## Choose the delivery form The endpoint can return audio in two forms, and the right choice depends on the consumer: - **Binary audio (`raw: true`)** — the response body is the audio file itself, served inline with its MIME type. Use this when a player, browser `