Source: http://localhost:1313/docs/getting-started.html

# Getting Started

This guide takes you from an AIVAX account to a verified OpenAI-compatible chat completion. The example uses Python and a private API key from a server-side environment.

By the end, you will have confirmed that your key and selected model or AI Gateway can complete a request.

## Before you begin

You need:

- An AIVAX account with dashboard access and permission to create a private API key.
- Python 3.8 or later with `pip` available.

For pricing and operational limits, see [Pricing](http://localhost:1313/docs/pricing.md) and [Plans and limits](http://localhost:1313/docs/limits.md).

Production API base URL:

```text
https://inference.aivax.net
```

OpenAI-compatible SDK base URL:

```text
https://inference.aivax.net/v1
```

## 1. Create a private API key

Create a **private** key from the API Keys area of the AIVAX dashboard. Copy the key when it is shown and store it as a secret; do not paste the real value into the code below.

Private keys are intended for trusted server-side applications. Public keys are restricted credentials for intentionally exposed client-side routes and are not a substitute for a backend key.

If you are building a public web widget or messaging experience, review [Chat clients](http://localhost:1313/docs/features/chat-clients.md) before exposing any credential. Chat sessions provide a clearer boundary for user identity, conversation history, and attachments.

See [Authentication](http://localhost:1313/docs/authentication.md) for supported authentication schemes, private and public key behavior, and secret-handling guidance.

## 2. Install the OpenAI SDK

Install the SDK in the Python environment you will use for this example:

```bash
python -m pip install openai
```

Keep the key outside your source file. For example, set an environment variable named `AIVAX_API_KEY` using the secret-management method appropriate for your shell or deployment platform.

## 3. Choose a model or AI Gateway

The `model` field can identify:

- A hosted model returned by the model listing endpoint.
- An AI Gateway available to your account.

Use a **hosted model** for a direct, one-off call or early experiment. Use an **AI Gateway** when you want to reuse the same model, instructions, RAG collections, skills, tools, moderation, and output settings across requests or users.

Gateway slugs are supported with private keys. Public-key chat completions must use the full gateway UUID and cannot call integrated models directly.

Use the model listing reference below to choose a hosted model. If you already have an AI Gateway, use its identifier instead.

[API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Model%20listing)

Copy one model name or gateway identifier that is available to your account. You will use it as `<MODEL_OR_GATEWAY_ID>` in the next step.

## 4. Make the first request

Create a file named `quickstart.py` with the following code:

```python
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://inference.aivax.net/v1",
    api_key=os.environ["AIVAX_API_KEY"],
)

response = client.chat.completions.create(
    model="<MODEL_OR_GATEWAY_ID>",
    messages=[
        {"role": "user", "content": "Write a one-sentence welcome message."}
    ],
)

print(response.choices[0].message.content)
```

Replace `<MODEL_OR_GATEWAY_ID>` with the exact hosted model name or gateway identifier selected in the previous step. Keep `AIVAX_API_KEY` unchanged in the code: it is the environment variable name, not the key value. Set that variable before running the file.

Run the file:

```bash
python quickstart.py
```

A successful request prints one generated sentence and exits without an API error.

Reference:

[API endpoint reference](https://inference.aivax.net/apidocs?embed=iframe&embed-endpoint=Inference%20(chat%20completions))

## 5. Confirm the integration

Confirm that the generated response matches the prompt and comes from the model or AI Gateway selected in the previous step. This verifies the endpoint, credential, and model selection used by your application.

Before increasing traffic or processing large inputs, review [Pricing](http://localhost:1313/docs/pricing.md) and [Plans and limits](http://localhost:1313/docs/limits.md).

## Troubleshoot the first request

AIVAX uses two response styles:

- OpenAI-compatible endpoints return an OpenAI-style `error` object.
- Account and administrative endpoints return an AIVAX response envelope with an error or a successful `data` value.

| Status | What to check |
| --- | --- |
| `400 Bad Request` | Confirm the model or gateway identifier and remove unsupported parameters from the request. |
| `401 Unauthorized` | Confirm that the private key is present, complete, active, and sent through the SDK configuration. |
| `402 Payment Required` | Review [Pricing](http://localhost:1313/docs/pricing.md) and confirm that the account is ready for a billable request. |
| `403 Forbidden` | Confirm that the key type, model, or selected resource allows this operation. |
| `429 Too Many Requests` | Retry later and review [Plans and limits](http://localhost:1313/docs/limits.md) before increasing request volume. |
| `500 Internal Server Error` | An unexpected AIVAX failure occurred. Retry later; the response does not include internal details. |
| `503 Service Unavailable` | A service AIVAX depends on is temporarily unavailable. Retry after the interval in the `Retry-After` header. |

If the request still fails, verify in this order:

1. `base_url` is `https://inference.aivax.net/v1`.
2. `AIVAX_API_KEY` is available to the Python process and contains a private key.
3. The selected model or gateway exists and is available to the account.
4. For a gateway, test a plain prompt before adding RAG, tools, media, or structured output so you can isolate configuration problems.

## Long-running inference

If a request ends with HTTP `524` or a proxy timeout while AIVAX is still processing it, use the direct inference host:

```text
https://direct.inference.aivax.net/v1
```

The request remains synchronous, not a background job: keep the client connection open until the completion finishes, and configure a client timeout that covers the expected processing time.

Use the same private API key, model or AI Gateway identifier, messages, and request parameters. Change the SDK base URL and timeout:

```python
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://direct.inference.aivax.net/v1",
    api_key=os.environ["AIVAX_API_KEY"],
    timeout=300.0,
)

response = client.chat.completions.create(
    model="<MODEL_OR_GATEWAY_ID>",
    messages=[
        {
            "role": "user",
            "content": "Analyze this case carefully and provide a detailed recommendation.",
        }
    ],
)

print(response.choices[0].message.content)
```

## Choose the next product

Once the minimal request works, add one capability at a time:

- [AI Gateways](http://localhost:1313/docs/inference/ai-gateway.md) — make the assistant configuration reusable across requests and users.
- [Structured responses](http://localhost:1313/docs/inference/structured-responses.md) — require generated JSON to follow an application schema.
- [RAG collections](http://localhost:1313/docs/rag/collections.md) — index your documents, test retrieval, and attach grounded knowledge to a gateway.
- [Built-in tools](http://localhost:1313/docs/tools/builtin-tools.md), [MCP](http://localhost:1313/docs/tools/mcp.md), or [Protocol functions](http://localhost:1313/docs/tools/protocol-functions.md) — let the assistant retrieve live information or take action.
- [Chat clients](http://localhost:1313/docs/features/chat-clients.md) — deliver a gateway through web chat or supported messaging channels.
- [Text and media products](http://localhost:1313/docs/overview.md#process-text-documents-and-media) — classify or segment documents, generate images or speech, transcribe audio, and describe media.
- [Batch](http://localhost:1313/docs/features/batch.md) — apply the same workflow to many independent records asynchronously.
- [Agentic Tests](http://localhost:1313/docs/inference/agentic-tests.md) — evaluate a complete gateway conversation before and after configuration changes.

Before increasing traffic or processing large inputs, review [Pricing](http://localhost:1313/docs/pricing.md) and [Plans and limits](http://localhost:1313/docs/limits.md).
