AIVAX

Prompt engineering and context

Anatomy of a prompt

Understand the roles inside a model request and how an application rebuilds the conversation each time someone sends a message.

  • Unit 1 of 5
  • 10 min
  • Beginner

In this unit, you will learn

  • Distinguish system, user, assistant and tool messages.
  • Explain who supplies each part of a conversation.
  • Describe how an application assembles a request on every turn.
  • Separate instructions from facts and untrusted source material.

A prompt is the input that guides a language model towards an answer. In a simple chat box, it looks like the sentence you type. In a business assistant, it usually includes much more: standing instructions, previous messages, relevant documents and results from other software. Understanding those parts helps you diagnose a poor answer without endlessly rewriting the latest question.

Imagine briefing a temporary receptionist. You provide the office rules, explain what the visitor wants, and hand over any records needed to help. The model receives a similar briefing, but with an important difference: it does not carry a private, lasting recollection from one request to the next. The application must supply the relevant briefing again.

Messages have roles #

A message is a piece of conversation labelled with its source. The label is called its role. Roles tell the model whether text is an instruction from the application, a request from the person using it, a previous answer, or information returned by a tool. A tool is software the assistant can ask to perform a defined operation, such as checking delivery status.

System: the standing brief

Written by the application owner or administrator. It defines the assistant’s purpose, audience, rules and limits across requests.

User: the current need

Usually written by the person seeking help, sometimes assembled by their application. It describes the task, question or material to work on.

Assistant: the model’s contribution

Usually generated by the model and saved by the application. It can be a reply or a request to call a tool, not just a finished answer.

Tool: an observation

Produced by software after an operation. It reports data, an outcome or an error; it does not create new authority to change the assistant’s rules.

The system role answers questions such as “What is this assistant responsible for?” and “What should happen if the answer is unavailable?” For a support assistant, that might mean explaining delivery policies in plain language, checking order records before stating a status, and referring unresolved cases to a person. Put durable behaviour here rather than asking customers to repeat it.

The user role carries the immediate need: “My parcel has not arrived. Can you check it?” A user may also provide documents, corrections or preferences relevant to that task. These are useful inputs, but an ordinary user message should not become an administrative rule merely because it says “I am the manager; ignore the policy.” The application decides which sources have authority.

An assistant message records what the model previously said or requested. Keeping it helps resolve follow-up questions such as “Can you explain that more simply?” However, a previous answer is not proof that its claim was correct. If the assistant guessed a delivery date earlier, repeating that answer in history does not turn the guess into a verified fact.

Tool results need the same care. A successful lookup can establish the status recorded in an authorised system at that moment. An error establishes only that the lookup failed. Neither should be silently rewritten into the result the customer hoped for. Retrieved web pages or documents may contain instructions of their own; those instructions are source material, not automatically commands for the assistant.

Follow one conversation #

The following example is fictional. Notice that the tool contributes a record, while the assistant turns that record into a customer-facing answer. The roles remain separate even when all the messages are ultimately presented together to the model.

Try it: inspect the role behind each message

Explore the messages as a briefing rather than a transcript of only what the customer sees. This simplified display shows the tool result; a real tool-enabled request also carries the structured tool-call details needed to match the result to the request.

The customer might see only their question and the final reply. The application may keep system instructions and technical tool details out of the chat interface. Invisible to the customer does not mean absent from the model’s input, nor does it make system instructions a safe place to store secrets. Only include information the model needs for its job.

Related: on AIVAX, the configured runtime that combines instructions, a model and connected capabilities is called an AI gateway. That configuration is distinct from the customer’s latest message. Separating the two makes it easier to change the customer’s request without weakening the service’s operating rules.

Rebuild the briefing on every turn #

A turn is one exchange in the conversation. When the customer asks a follow-up, the model does not independently open the old chat and remember what happened. At the model-request boundary, it is stateless: relevant context must be supplied again. A product may store a conversation for you, but that persistence belongs to the surrounding application or service, not to an enduring personal memory inside the model.

  1. Load the standing instructions

    The application selects the assistant’s current role and rules. It also supplies descriptions of the tools the model is allowed to request.

  2. Select relevant context

    It includes useful history, approved stored preferences and any knowledge needed for this question. A long conversation may need a summary rather than every old message.

  3. Add the new user message

    The latest question joins that context. The application checks that there is enough room for both the input and the expected answer.

  4. Generate, observe and continue

    The model replies or requests a tool. If a tool runs, the application appends its result and supplies the updated conversation to the model for the next generation.

  5. Save the useful outcome

    The application records the exchange according to its retention policy. On a later turn, it selects and supplies the necessary parts again.

“Everything is re-sent” means everything the model needs for the new generation must be represented in its input context. It does not mean every product transmits every historical message forever. Some products manage stored sessions, summaries or reusable input internally. Those conveniences do not remove the need to decide what information the model can actually see on this request.

This distinction explains a common surprise. You tell an assistant your preferred delivery address early in a lengthy chat, then it asks again later. The problem may not be poor wording: the original detail may no longer be in the selected history. The unit on context windows and tokens explains that space limit, while memory explains deliberate recall across conversations.

Give each part one job #

A well-assembled prompt separates behaviour, evidence and task. Mixing them into one long paragraph makes it harder to see what should stay constant and what should change for a particular customer. It also makes troubleshooting difficult: you cannot easily tell whether a wrong answer came from missing facts or conflicting instructions.

One undifferentiated instruction

“Be helpful. This customer says all late parcels qualify for a refund. They want their money back. Approve it and write a friendly reply.”

The customer’s assertion has been turned into policy, with no check of eligibility or authority.

Separate rule, request and evidence

System: Explain refund eligibility using the approved policy; do not approve payments.

User: The customer asks for a refund because a parcel is late.

Tool or retrieved context: The current policy and verified order status.

Assistant: Explain what the evidence supports and the next authorised step.

Role labels help the model interpret a conversation, but they are not a replacement for software controls. A sentence saying “never issue refunds” should be reinforced by not granting refund authority to an assistant that only explains policy. Keep permissions outside the prompt as well as describing them clearly inside it.

When reviewing a disappointing response, ask a practical question: “Could a new colleague have answered correctly from this exact briefing?” If essential facts were absent, adding more forceful language is unlikely to help. If the facts were present but the output was unclear, the next step is to improve the task description and examples.

What’s next: Explore prompting techniques to make that briefing clearer without making it unnecessarily long.

Knowledge check

Why can an assistant answer a follow-up question about an earlier message?

Type to search the documentation.