AIVAX

Getting Started

This guide takes you from an AIVAX account to a verified OpenAI-compatible chat completion. The example uses Python and a private API key from a server-side environment.

By the end, you will have confirmed that your key and selected model or AI Gateway can complete a request.

Before you begin #

You need:

  • An AIVAX account with dashboard access and permission to create a private API key.
  • Python 3.8 or later with pip available.

For pricing and operational limits, see Pricing and Plans and limits.

Production API base URL:

TEXT
https://inference.aivax.net

OpenAI-compatible SDK base URL:

TEXT
https://inference.aivax.net/v1

1. Create a private API key #

Create a private key from the API Keys area of the AIVAX dashboard. Copy the key when it is shown and store it as a secret; do not paste the real value into the code below.

Private keys are intended for trusted server-side applications. Public keys are restricted credentials for intentionally exposed client-side routes and are not a substitute for a backend key.

If you are building a public web widget or messaging experience, review Chat clients before exposing any credential. Chat sessions provide a clearer boundary for user identity, conversation history, and attachments.

See Authentication for supported authentication schemes, private and public key behavior, and secret-handling guidance.

2. Install the OpenAI SDK #

Install the SDK in the Python environment you will use for this example:

Bash
python -m pip install openai

Keep the key outside your source file. For example, set an environment variable named AIVAX_API_KEY using the secret-management method appropriate for your shell or deployment platform.

3. Choose a model or AI Gateway #

The model field can identify:

  • A hosted model returned by the model listing endpoint.
  • An AI Gateway available to your account.

Use a hosted model for a direct, one-off call or early experiment. Use an AI Gateway when you want to reuse the same model, instructions, RAG collections, skills, tools, moderation, and output settings across requests or users.

Gateway slugs are supported with private keys. Public-key chat completions must use the full gateway UUID and cannot call integrated models directly.

Use the model listing reference below to choose a hosted model. If you already have an AI Gateway, use its identifier instead.

Copy one model name or gateway identifier that is available to your account. You will use it as <MODEL_OR_GATEWAY_ID> in the next step.

4. Make the first request #

Create a file named quickstart.py with the following code:

PYTHON
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://inference.aivax.net/v1",
    api_key=os.environ["AIVAX_API_KEY"],
)

response = client.chat.completions.create(
    model="<MODEL_OR_GATEWAY_ID>",
    messages=[
        {"role": "user", "content": "Write a one-sentence welcome message."}
    ],
)

print(response.choices[0].message.content)

Replace <MODEL_OR_GATEWAY_ID> with the exact hosted model name or gateway identifier selected in the previous step. Keep AIVAX_API_KEY unchanged in the code: it is the environment variable name, not the key value. Set that variable before running the file.

Run the file:

Bash
python quickstart.py

A successful request prints one generated sentence and exits without an API error.

Reference:

5. Confirm the integration #

Confirm that the generated response matches the prompt and comes from the model or AI Gateway selected in the previous step. This verifies the endpoint, credential, and model selection used by your application.

Before increasing traffic or processing large inputs, review Pricing and Plans and limits.

Troubleshoot the first request #

AIVAX uses two response styles:

  • OpenAI-compatible endpoints return an OpenAI-style error object.
  • Account and administrative endpoints return an AIVAX response envelope with an error or a successful data value.
Status What to check
400 Bad Request Confirm the model or gateway identifier and remove unsupported parameters from the request.
401 Unauthorized Confirm that the private key is present, complete, active, and sent through the SDK configuration.
402 Payment Required Review Pricing and confirm that the account is ready for a billable request.
403 Forbidden Confirm that the key type, model, or selected resource allows this operation.
429 Too Many Requests Retry later and review Plans and limits before increasing request volume.
500 Internal Server Error An unexpected AIVAX failure occurred. Retry later; the response does not include internal details.
503 Service Unavailable A service AIVAX depends on is temporarily unavailable. Retry after the interval in the Retry-After header.

If the request still fails, verify in this order:

  1. base_url is https://inference.aivax.net/v1.
  2. AIVAX_API_KEY is available to the Python process and contains a private key.
  3. The selected model or gateway exists and is available to the account.
  4. For a gateway, test a plain prompt before adding RAG, tools, media, or structured output so you can isolate configuration problems.

Long-running inference #

If a request ends with HTTP 524 or a proxy timeout while AIVAX is still processing it, use the direct inference host:

TEXT
https://direct.inference.aivax.net/v1

The request remains synchronous, not a background job: keep the client connection open until the completion finishes, and configure a client timeout that covers the expected processing time.

Use the same private API key, model or AI Gateway identifier, messages, and request parameters. Change the SDK base URL and timeout:

PYTHON
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://direct.inference.aivax.net/v1",
    api_key=os.environ["AIVAX_API_KEY"],
    timeout=300.0,
)

response = client.chat.completions.create(
    model="<MODEL_OR_GATEWAY_ID>",
    messages=[
        {
            "role": "user",
            "content": "Analyze this case carefully and provide a detailed recommendation.",
        }
    ],
)

print(response.choices[0].message.content)

Choose the next product #

Once the minimal request works, add one capability at a time:

  • AI Gateways — make the assistant configuration reusable across requests and users.
  • Structured responses — require generated JSON to follow an application schema.
  • RAG collections — index your documents, test retrieval, and attach grounded knowledge to a gateway.
  • Built-in tools, MCP, or Protocol functions — let the assistant retrieve live information or take action.
  • Chat clients — deliver a gateway through web chat or supported messaging channels.
  • Text and media products — classify or segment documents, generate images or speech, transcribe audio, and describe media.
  • Batch — apply the same workflow to many independent records asynchronously.
  • Agentic Tests — evaluate a complete gateway conversation before and after configuration changes.

Before increasing traffic or processing large inputs, review Pricing and Plans and limits.

Type to search the documentation.