
The popularity of agents is surging. Many have claimed that LLM-backed applications or agents can bridge the gap between prompt and production faster than you can say AI. Yet, with new AI models replacing each other every week and AI solutions promising rapid productivity boosts, investors, technology and business leaders still find themselves asking the same practical questions:
By now, you must have read the signposts framing the agentic landscape. Agents are a new way of organizing and delegating work to AI-backed systems which perform one or more functions (connected to your software, data, and other digital tools) within set boundaries (rules, limits, and checks). However, agents do not run themselves forever. Even after being set up intelligently, they require maintenance and supervision from a HITL (human-in-the-loop), other specialized AI-systems, or a combination of both. Building agents takes many rounds of design, testing and adjustment before production, as well as reinforcement throughout the agent lifecycle.
While many AI advocates claim that anyone can build smart, capable agents, the reality on the ground has already shown that, unless you understand your business needs, have standardized processes, and have deep domain expertise on your team, the agents you will build will be short-lived, buggy, and inefficient.
Poorly built agentic systems have already brought an uncomfortable dent into businesses: public shame and exposure, ruffled reputation, disgruntled customers, and churning revenue. Air Canada was ordered by a tribunal to pay back a discount its chatbot invented, and Cursor lost paying customers after its own support bot invented a company policy that never existed. Agents may have become popular, but more so have the businesses implementing faulty agents.
Every time agents pass wrong information as the truth, they are said to hallucinate. Agents can fabricate what is not there or simply mislead you by agreeing with your faulty premises.
Give a home cook and a chef the same chicken, and only one knows not to put it frozen into a scorching oven. Agents are no different: they are only as good as the people who build them and the checks around them.
Agents can be built fast, but they must be built with a future-proof vision in mind. Agents need to be fast, reliable, and relevant today, next week, during your annual board meeting, and every single day. Agents require governance. Agents need AI evals (evaluations).
Let’s start with a story which did not begin once upon a time, but in our time, in your very back (business) yard: The dashboard that lied. You built an AI-based IT helpdesk automation and called it Sam. Among its many functions, Sam is in charge of solving IT tickets whenever an employee account needs to be activated or deactivated. Sam does a fine job. The average ticket resolution time improves every single day. The dashboard is green. Sam must be doing a good job, you would say. However, a very different story unfolds behind the scenes. Sam quietly enables access for accounts that should have been deactivated. Sam sees both the HR record and the identity system record, but treats an account as active if either system says so. Recently departed employees get their access refreshed because they are still in the identity system, as HR offboarding and IT deprovisioning are out of sync. Good old Sam was not instructed to check for contradictions. Tickets are closed successfully, service-level targets are met, the metrics are booming. An external audit compares resolved tickets against HR employment records. First comes the shock, then the realization. A fine is issued, audit score drops, internal investigation identifies the culprit: Sam has some pieces missing and there is no mechanism that checks the calls it makes, the AI tools it selects, and its own self-correction attempts. Sam was never designed with the proper, embedded guardrails in place, or at least not in production. Sam is missing AI evals and is failing silently.
So, what are AI evals? AI evals are tests that check whether your AI system is doing what it is supposed to be doing. They come in two kinds:
The AI models behind agents do not always give the same answer to the same request, and a change in the model, the instructions, or the data can make an agent quietly worse at a task it used to handle well. Offline evals catch this before release. Online evals catch it after. Sam had neither. Offline evals would have flagged the missing HR check before launch. Online evals would have flagged the first reactivated ex-employee account. Instead, the dashboard stayed green, Sam carried on, and so did you.
AI evals are needed both before building your agent and after. But how do you get started?
Your AI evals are as proficient as the level of design and sophistication invested before production and after production deployment.
So what does this mean? Let’s look at an analogy first. You need to sit an important test. You go through the handbook and practice tests. You are very accustomed to the test format and you do study diligently. You pass the test with flying colours. Then, you get a job in the field. The job does not have a test format. You need to know how to apply the handbook on the fly. You realize you are not that prepared for the real world. Being a good student and passing tests whose format is familiar is not the same as solving real-life issues.
In the same vein, creating proficient AI evals that are objective enough, smart enough, and flexible enough must be anchored in various good practices which must complement any skilled development team:
Domain experts and AI teams work in tight collaboration to create the eval datasets. The people who know what a correct result looks like (your domain experts in operations, HR, finance, or compliance) define the test cases and review failures. In Sam’s case, that would have been HR and IT security, together. The good news is that AI evals are a professional skill which can be taught, nurtured and applied successfully.
Building reliable AI agents requires a heavier initial investment: the right platform, the right AI eval cases, the right labels. But consider the alternative. You ship an AI agent in a matter of weeks and pay a hefty production cost after a couple of weeks of interaction.
Without AI evals in place failures will occur. But will the failure be visible or invisible? Will you notice that the system is failing your business when it is too late?
With AI evals in place failures will occur, but most will be caught early, isolated safely, and resolved behind the scenes.
AI evals are a prerequisite for any agentic implementation. Here is what they give you:
Deloitte refunded the Australian government more than A$97,000 (about US$63,000) of an A$440,000 contract after a report shipped with a fabricated court quote and references to research that does not exist. Alphabet lost about $100 billion in market value in a single day after Bard’s launch demo gave a wrong answer about the James Webb Space Telescope. Next to figures like these, the cost of running evals is a rounding error.
If you are investing in a company that runs AI agents, the same logic applies to due diligence. An agent without evals is an unpriced risk: it can look productive on a dashboard while building up liabilities the numbers do not show. Ask to see eval results over time, not just a demo. Our FAQ below lists the questions to start with.
Before your next agent goes live, or your next investment closes, ask how it is tested. If the answer is a green dashboard, talk to us. Bonsai Labs helps teams design, build, and run AI evals that catch problems before customers, auditors, or investors do.
At Bonsai Labs, we start where the problems are, not where the tools are. Before we write a single eval, we study how your workflows fail today: the exceptions, the edge cases, and the contradictions your teams still handle by hand. We build evals from those real failures, not from generic examples, and combine fixed rules, AI judges, and review by your domain experts. Every eval becomes part of your release process, so a change that lowers quality does not reach production. Once the agent is live, we monitor its work and turn every new failure into a new test. The result is an agent whose performance you can see, measure, and explain to your customers, your auditors, and your board.
An AI eval is a test that checks whether your AI system does what it is supposed to do. Some evals check hard facts, such as whether the agent picked the right tool or returned a valid record. Others judge quality, such as whether an answer is accurate, complete, and in line with your policies.
Traditional software gives the same output for the same input, so a test passes or fails the same way every time. AI models can answer the same request differently, and a change in the model, the instructions, or the data can quietly shift quality. That is why evals run on many realistic cases and keep running after release.
Offline evals run before release against a fixed set of test cases, and can block a release if scores drop. Online evals run in production as part of monitoring and flag when the agent’s live work starts to slip. You need both: one keeps problems out, the other catches the ones that get through.
Engineering builds and runs them. Domain experts, the people who know what a correct answer looks like, define the test cases and the pass criteria. Technology and business leaders set the quality thresholds and decide how much risk the organization will accept.
Expect a meaningful share of an agent project to go into understanding failures and building evals, especially at the start. Running evals is comparatively cheap: engineering time to build the test set and harness once, model costs that are usually small per run, and possibly a platform subscription. Compare that with the cost of a single public failure.
There are many off-the-shelf platforms that give you the infrastructure to host and run your evals, such as Langfuse, Langsmith, Braintrust.
LLM-as-a-Judge means using a second AI model to grade your agent’s answers against a clear rubric, for qualities fixed rules cannot check, such as accuracy, tone, or whether instructions were followed. You can rely on it once its grades consistently agree with those of your human experts on the same sample. Judges have known biases, such as favoring longer answers, so a human should still review a sample regularly.
Run offline evals every time something changes: the prompt, the model, the tools, or the data the agent relies on. Run online evals continuously, or on a set schedule, for as long as the agent is in production. Add every new production failure to your test set so the same problem cannot return unnoticed.
Good evals are built from your real failures, not generic examples, and they check each step of the agent’s work as well as the end result. Their verdicts agree with those of your human experts. If production keeps surprising you, your evals have gaps.
Ask what the company tests before each release, and whether a drop in eval scores can block a release. Ask how it monitors the agent in production, who reviews failures, and how those failures feed back into the tests. A team that can show its eval results over time understands its AI; a team that can only show a green dashboard may not.