Semantic decisions
Semantic decisions evaluate named questions against a shared state and return structured answers rather than a generated explanation. Use them to route support requests, select a category, check a condition, or assign an ordered score.
A request supplies the model, the evidence in state, and a questions object. Each question has an ID you choose; the response uses the same ID in answers. You can ask different question types in one request without making separate API calls.
Use structured responses instead when you need a larger generated object or a written explanation. For embedding-based similarity between documents and labels, see Text classification.
Choose a model
All models below support choice, noul, and score. Prices are base USD prices per million input tokens; account and plan adjustments still apply. Output tokens have no charge in the current decision-model catalog.
| Model | Input price / million tokens | Context |
|---|---|---|
@supersonic-labs/julia-1 |
$0.008 | 1,024 tokens per question |
@typesafe/jev-1.13 |
$0.042 | 32,768 tokens |
@respan/span-01 |
$0.020 | Not specified in the current catalog |
@respan/span-01-lite |
$0.000 | Not specified in the current catalog |
@jaredpalmer/kev-4b |
$0.042 | 8,192 tokens |
@typesafe/jev is also accepted and currently resolves to @typesafe/jev-1.13. An unspecified context limit does not mean unlimited input. Model-specific limits and interpretation of scores can differ; validate a model on representative examples before switching production traffic.
An authenticated API key and a positive account balance are required, including when selecting a model with a zero base token price. See Authentication, Pricing, and Plans and limits.
Discover models programmatically
GET /api/v1/information/decisions-models.json lists the current decision-model catalog without authentication. The response's data array contains:
name: the canonical identifier to use in a decision request.aliases: other accepted identifiers for that model.contextLength: the advertised context in tokens, ornullwhen unspecified.releaseDate: the catalog release date inyyyy-MM-ddformat.capabilities: supported question types (noul,choice, and/orscore).inputPricePerMillionTokensandoutputPricePerMillionTokens: base USD prices, before account and plan adjustments.
Use this listing to populate model selectors rather than maintaining a separate hardcoded catalog. It describes configured models, not a real-time provider health check. Model-specific option and question limits are not included in this listing.
Write the questions
| Type | criteria format |
Use it for |
|---|---|---|
choice |
Object mapping choice IDs to descriptions | Select one destination or category |
noul |
Object with nonempty false and true descriptions |
Evaluate a Boolean condition |
score |
Ordered array of level descriptions | Evaluate a graded condition |
Every question requires nonempty instructions. Use descriptions that distinguish the options, not just opaque IDs. For example, "billing": "Charges, payments, and refunds" provides more evidence than "billing": "B".
For noul, both criteria are required. Describe what counts as false as carefully as what counts as true, especially when the state might omit the relevant information. For score, keep the level order consistent across requests.
Evaluate a support request
Send the following JSON body to POST /api/v1/generations/decisions. The example uses Julia-1; its state may be text, an object, or an array.
{
"model": "@supersonic-labs/julia-1",
"state": {
"message": "I was charged twice. Please return the extra payment."
},
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Charges, payments, and refunds",
"technical": "Software errors and service outages",
"sales": "Plans, pricing, and purchases"
}
},
"refund_requested": {
"type": "noul",
"instructions": "Does the customer explicitly ask for money back?",
"criteria": {
"false": "The customer does not ask for money to be returned",
"true": "The customer asks for a refund or return of a payment"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is the request based on the stated deadline?",
"criteria": [
"No deadline stated",
"A deadline is stated, but it is not today",
"The customer explicitly needs resolution today"
]
}
}
}
Keep only relevant evidence in the state. Instructions should explain the decision, not ask for a chain of reasoning or additional text. Avoid overlapping choice descriptions unless that ambiguity is intentional.
Read the response
The success response is a JSON object directly, without a data envelope. It contains:
id: the decision identifier.model: the canonical model identifier used for the request.provider: the provider reported for the result.answers: an object keyed by your question IDs.usage:input_tokens,output_tokens, and the billedcost.
For Julia-1, an illustrative answers object for the example above is shown below. These numbers explain the format; they are not a recorded response or a quality guarantee.
{
"department": {
"type": "choice",
"choice": "billing",
"probabilities": {
"billing": 0.90,
"technical": 0.06,
"sales": 0.04
}
},
"refund_requested": {
"type": "noul",
"noul": 0.95,
"probabilities": {
"false": 0.05,
"true": 0.95
}
},
"urgency": {
"type": "score",
"score": 0.3,
"legend": {
"0": "No deadline stated",
"1": "A deadline is stated, but it is not today",
"2": "The customer explicitly needs resolution today"
},
"probabilities": {
"0": 0.8,
"1": 0.1,
"2": 0.1
}
}
}
With Julia-1:
choiceis the selected caller-defined ID, not the option description.noulis the probability assigned to the true criterion, not a JSON Boolean. Your application chooses the threshold and how to handle uncertain cases.scoreis the expected zero-based level index. In the example,0 × 0.8 + 1 × 0.1 + 2 × 0.1 = 0.3. It is not necessarily an integer, and it is not a normalized 0–1 score when there are more than two levels.probabilitiesare keyed by choice IDs,false/true, or score-level indices.legenddescribes the score levels.
Other models may omit optional fields such as probabilities, legend, or confidence. Do not assume that every provider uses the same score scale or confidence definition. A high probability is not proof that the decision is correct; validate thresholds and escalation rules with labeled examples from your own domain.
Account rate limits
Semantic decision requests share an account-level rate limit across models and API keys. Each request counts once, even when it contains multiple questions. This quota is separate from the daily subscription allowance and applies to both included and paid usage. See Plans and Limits for the Free, Pro, and Max thresholds.
Requests above the limit return 429 Too Many Requests before evaluation. Pace calls across the account and retry with backoff after the rate-limit window clears; changing API keys within the same account does not provide a separate quota.
Julia-1 limits
| Limit | Value |
|---|---|
| Questions per request | 1–32 |
| Choices or score levels per question | 2–20 |
| Boolean options | Exactly two: false and true |
| Combined context per question | 1,024 tokens, including state, question, options, and special tokens |
| Question and options budget | 256 tokens within the combined context |
| Description of an individual option | At most 48 tokens |
| Decision payload limit | 256 KiB |
Limits interact: twenty options can exceed the combined question/options budget even if every description is individually short enough. Shorten descriptions or reduce the option count instead of assuming every maximum can be used at once.
Inputs exceeding the context or question/options budget are rejected, not silently truncated. The literal <mask> is reserved and cannot appear in the state, instructions, or option descriptions. These are the current AIVAX Julia-1 serving limits; larger context figures in an upstream model card do not override them.
Usage and cost
For Julia-1, input usage sums the encoded sequence for each question, excluding padding. The shared state is therefore counted again for each question. Four questions over one state do not have the same input usage as one question over that state. Julia-1 does not generate text, so its output_tokens value is zero.
Julia-1 is currently eligible for the daily semantic decision allowance on Free, Pro, and Max. Other decision models are billed normally. The allowance is shared across eligible decision calls, not reserved for each question or API key. See Plans and Limits for relative plan capacity and coverage rules.
When not covered, Julia-1 input is billed at the base price of $0.008 per million tokens before account and plan adjustments. Use the returned usage.cost for the actual billed amount; it is zero when the input is fully covered by the allowance.
Errors and reliable use
- Invalid model or question: check the exact model identifier, question type, instructions, and criteria shape. Question IDs and choice IDs must be nonempty.
- Context or option limit exceeded: shorten the state or descriptions, reduce the option count, or select a model with suitable limits. Retrying the same invalid input will not resolve it.
- Authentication or balance error: verify the API key and account balance before retrying. A zero-priced model still requires a positive balance.
- Rate limit (429): reduce the account's request rate and retry with backoff. Multiple questions in one request still count as one request, but model-specific payload limits and per-question usage remain applicable.
- Temporary capacity or provider unavailability: avoid an immediate parallel retry storm. Reduce concurrency and use bounded retries with backoff for transient failures.
Evaluate choice, noul, and score separately when validating a model: success at routing does not establish reliable scoring or Boolean behavior. Include ambiguous and incomplete states in your test set, and use human review where a wrong decision has material consequences. A retry is a new request; do not assume automatic deduplication or identical model outputs.
English
Português