Elevator panel with Context, Loop, Jev, Harness and Evals buttons; the Harness button is illuminated

A working AI agent receives a task, retrieves information, uses tools and adjusts its actions based on the results. The language model interprets the request and selects the next step. To complete the task, it needs a software environment that supplies information, executes operations and checks the outcome.

An AI agent's architecture can be understood through five engineering tasks: preparing context, managing the action loop, routing requests, providing an execution environment and evaluating quality. This division helps identify what the system consists of and which part needs attention when a problem occurs.

The layers work together. Their numbering sets the order of this explanation; the implementation depends on the task. Jev is a specific TypeSafe product. In the broader architecture, that section covers request classification and routing.

Layer Responsibility
Context What information the model receives before its next action
Loop How the next step is selected and when the work ends
Jev / routing Which decisions can be handled separately from the main agent
Harness Where the agent operates and what it is allowed to do
Evals How quality is measured and regressions are detected

1. Context engineering: preparing information for the model

Context includes instructions, conversation history, retrieved documents, tool descriptions and the results of previous operations. Its contents determine which information the model can use when choosing its next action.

The size of the context window alone does not determine response quality. A long history may contain useful information alongside outdated conclusions, repeated messages and large technical dumps. Context engineering therefore involves selecting the relevant information at each step.

In practice, this means making several design decisions:

  • Stable instructions define the role, constraints and requirements for the result.
  • Documents and large computation results remain in external sources. The model retrieves the relevant portion through an available tool.
  • The history preserves important decisions, the current objective and unresolved questions.
  • Tools are selected for the task. A large catalogue can be searched for the appropriate tool instead of loading every description at once.

A document link is useful only if the agent can actually open it. Similarly, a summary of the history must retain the constraints and unresolved questions that will matter later.

Reusing the unchanged part of a request through prompt caching is a separate consideration. In the Claude API, for example, a cache hit requires an exact match of the relevant prefix. It therefore makes sense to separate stable instructions from changing data. Savings depend on the provider's terms and the actual cache hit rate.

For a coding agent, useful context includes the bug description, relevant files, project rules and test results. The entire repository history is usually unnecessary for each model call.

2. Loop engineering: managing the sequence of actions

In a predefined workflow, the program determines the order of steps. In an agent loop, the model chooses the next step after receiving the result of the previous one. Anthropic's architecture guidance explains this distinction in detail.

When fixing a bug, an agent might run tests, examine the failure, open the relevant file, make a change and run the checks again. The sequence depends on the cause it discovers. The conditions for considering the fix complete are defined in advance.

A controlled loop needs four elements: a specific objective, verification of the result, stopping conditions and a budget. Verification may be a test, a successful build, a data reconciliation or a confirmed state in an external system. The model saying “done” is not sufficient verification.

Stopping conditions cover successful completion, lack of progress and reaching a limit. The budget caps the number of actions, elapsed time and cost. An unfinished task needs a defined next step: clarification, a fixed procedure or a handover to a member of staff.

An agent loop is useful when the course of work depends on intermediate results. A known sequence of operations is often best kept as a conventional workflow. Both approaches can coexist in one system: standard cases follow predefined rules, while the agent handles exceptions.

3. Jev engineering: classification and routing

Some system tasks involve narrow decisions: identifying the category of a request, choosing a handler or scoring a message against a defined criterion. These operations can form a separate layer before the main agent starts.

Jev from TypeSafe is one tool for these decisions. According to its documentation, the model accepts data and questions with a predefined response format. It supports choosing from a list, scoring on a scale and assigning a numerical estimate to the truth of a statement. The result consists of structured values that the program can use.

For example, incoming requests can be routed to documentation, routine question handling or technical investigation. Only the last route requires an agent to investigate the problem step by step and select its actions.

A valid response format does not guarantee correct classification. A selected category can be permitted by the schema and still be wrong. In the Jev announcement, the developer specifically ties the absence of formatting errors to schema compliance.

Confidence scores also need validation on the company's own data. Jev derives them from the probability distribution for Choice and Score; Noul has no separate confidence score. Routing thresholds should reflect measured quality and the consequences of an error.

Before deployment, compare this layer with simple rules or a conventional classifier. Uncertain and unusual requests should receive closer investigation. Jev is one implementation option; results on the company's tasks should determine the choice of tool. A separate Journal article examines the product in more detail.

4. Harness engineering: the execution environment and access control

The harness is the software environment around the model: tools, the file system, project rules, isolation, access permissions and operation checks. It determines what the agent can actually do.

This environment has four main responsibilities.

Isolation. The agent works in a dedicated environment with limited access to files, networks and data. For development, this might be a container with a separate working copy of the project and access only to a test database.

Instructions. Project rules, tool descriptions and examples of the expected result guide the choice of actions. In a repository, these rules can be stored in AGENTS.md.

Checks. Builds, tests, type checking and other automated checks provide feedback on the work. An operation log helps establish what happened before an error.

Permissions. Access is limited to the operations required. Preparing a change, publishing it and modifying production data may require different permissions.

The execution system enforces these limits. An instruction in a prompt does not revoke write access granted to an account. OWASP recommends minimising available functions and permissions, enforcing authorisation outside the model and obtaining approval for actions with significant consequences.

The environment needs reviewing when the model or its tasks change. New capabilities may simplify parts of the control logic, but any change to the restrictions should be supported by testing.

5. Evals engineering: measuring quality

Evals are a repeatable set of tasks and assessment criteria. They make it possible to compare agent versions after changing the model, instructions, tools or routing rules.

Checks cover both the outcome and significant actions. For a coding agent, the outcome is a fixed bug with the required behaviour of the application preserved. Separate checks can verify compliance with access boundaries and mandatory approval before publication. Avoid prescribing the exact path unnecessarily: different valid solutions may use different sequences of steps.

The initial evaluation set should come from real tasks and observed failures. Each case needs inputs, an expected result and a verification method. Add cases as new scenarios emerge.

Objective criteria can be checked programmatically. Another model can help assess meaning, provided its judgements are first compared with those of specialists. Anthropic's agent evaluation methodology describes this approach.

Review results by category: an overall score can hide deterioration in a critical scenario. Alongside quality, track completion time, cost per task and the need for staff intervention.

How the five layers work together

Consider a hypothetical task: fixing a bug in a web application. Routing identifies the request type and passes it to a coding agent. Context supplies the bug report, project rules and relevant files. The action loop connects diagnosis, changes and repeated testing. The harness provides the working environment, tools and access boundaries. Evals check whether a new agent version maintains its quality across a set of similar tasks.

A ready-made platform may provide some of these capabilities. When choosing a solution, assess how well it fits the company's tasks: which data sources and integrations it supports, how it limits the agent's permissions and how it allows results to be evaluated.

This structure helps locate the source of a problem. Missing information calls for work on context. Repeated actions without progress call for a review of the loop. Incorrect request allocation calls for routing checks. Excessive access requires changes to the execution environment. Comparing evaluation results reveals quality changes after an update.

To discuss an AI agent's architecture with LindenTech, prepare one practical task, the data sources involved, the permitted actions and an example of an acceptable result. These provide a basis for determining which parts of the system are already available and what needs to be developed.