Bonsai Labs
SoftwareServices
BlogCase Studies
CareersGet in Touch
SoftwareServices
BlogCase Studies
CareersGet in Touch

© 2026 Bonsai Labs. All rights reserved.

Bonsai Labs OÜ, 17403676, Sepapaja tn 6, 15551 Tallinn, Harju Maakond, Estonia

EventsPrivacy PolicyCookie Policy
←All articles

Is your AI agent really working? What executives need to know about AI evals

October 9, 2026·Csongor Barabasi
Evals & Monitoring for AI Agents

The popularity of agents is surging. Many have claimed that LLM-backed applications or agents can bridge the gap between prompt and production faster than you can say AI. Yet, with new AI models replacing each other every week and AI solutions promising rapid productivity boosts, investors, technology and business leaders still find themselves asking the same practical questions:

  • How do I know my AI agent is still doing its job after launch?
  • How much human oversight does an agent need once it runs on its own?
  • What does a green dashboard actually tell me about my agent’s performance?
  • Who is accountable when an agent makes the wrong call?
  • What should I ask a vendor or a portfolio company about how they test their AI?

By now, you must have read the signposts framing the agentic landscape. Agents are a new way of organizing and delegating work to AI-backed systems which perform one or more functions (connected to your software, data, and other digital tools) within set boundaries (rules, limits, and checks). However, agents do not run themselves forever. Even after being set up intelligently, they require maintenance and supervision from a HITL (human-in-the-loop), other specialized AI-systems, or a combination of both. Building agents takes many rounds of design, testing and adjustment before production, as well as reinforcement throughout the agent lifecycle.

While many AI advocates claim that anyone can build smart, capable agents, the reality on the ground has already shown that, unless you understand your business needs, have standardized processes, and have deep domain expertise on your team, the agents you will build will be short-lived, buggy, and inefficient.

Poorly built agentic systems have already brought an uncomfortable dent into businesses: public shame and exposure, ruffled reputation, disgruntled customers, and churning revenue. Air Canada was ordered by a tribunal to pay back a discount its chatbot invented, and Cursor lost paying customers after its own support bot invented a company policy that never existed. Agents may have become popular, but more so have the businesses implementing faulty agents.

Every time agents pass wrong information as the truth, they are said to hallucinate. Agents can fabricate what is not there or simply mislead you by agreeing with your faulty premises.

Give a home cook and a chef the same chicken, and only one knows not to put it frozen into a scorching oven. Agents are no different: they are only as good as the people who build them and the checks around them.

Agents can be built fast, but they must be built with a future-proof vision in mind. Agents need to be fast, reliable, and relevant today, next week, during your annual board meeting, and every single day. Agents require governance. Agents need AI evals (evaluations).

Are AI evals a necessity or hype?

Let’s start with a story which did not begin once upon a time, but in our time, in your very back (business) yard: The dashboard that lied. You built an AI-based IT helpdesk automation and called it Sam. Among its many functions, Sam is in charge of solving IT tickets whenever an employee account needs to be activated or deactivated. Sam does a fine job. The average ticket resolution time improves every single day. The dashboard is green. Sam must be doing a good job, you would say. However, a very different story unfolds behind the scenes. Sam quietly enables access for accounts that should have been deactivated. Sam sees both the HR record and the identity system record, but treats an account as active if either system says so. Recently departed employees get their access refreshed because they are still in the identity system, as HR offboarding and IT deprovisioning are out of sync. Good old Sam was not instructed to check for contradictions. Tickets are closed successfully, service-level targets are met, the metrics are booming. An external audit compares resolved tickets against HR employment records. First comes the shock, then the realization. A fine is issued, audit score drops, internal investigation identifies the culprit: Sam has some pieces missing and there is no mechanism that checks the calls it makes, the AI tools it selects, and its own self-correction attempts. Sam was never designed with the proper, embedded guardrails in place, or at least not in production. Sam is missing AI evals and is failing silently.

So, what are AI evals? AI evals are tests that check whether your AI system is doing what it is supposed to be doing. They come in two kinds:

  • Offline evals run before release. You test your agent against a set of cases, confirm it meets your standard, and release with confidence.
  • Online evals run in production, as part of monitoring. They score the agent’s live work and alert you when quality starts to slip.

The AI models behind agents do not always give the same answer to the same request, and a change in the model, the instructions, or the data can make an agent quietly worse at a task it used to handle well. Offline evals catch this before release. Online evals catch it after. Sam had neither. Offline evals would have flagged the missing HR check before launch. Online evals would have flagged the first reactivated ex-employee account. Instead, the dashboard stayed green, Sam carried on, and so did you.

Are AI evals needed before or after building agents?

AI evals are needed both before building your agent and after. But how do you get started?

  1. Understand your business: workflows, applications, errors, pain points. Do not automate for the sake of automating. Automate where it matters.
  2. Observe your production issues, their nature, the very things you want an agent to fix or prevent. In other words, understand your production errors and do not spare the time spent on analyzing what does not work. Some agent practitioners report spending 60–80% of their development time on error analysis and evaluation.
  3. Write your AI evals for pre-production environments. Design AI evals as binary (Pass/Fail) assessments of system checkpoints. If you work in a regulated industry, do not stop at anonymized test data. Many teams test on anonymized data, see everything pass, and then watch the agent break in production, where real records contain personal (PII) or health (PHI) information, messier formats, and edge cases the test set never had. Run your evals on production-grade data too, or on data as close to it as your compliance rules allow. Sam’s test data never contained an ex-employee still listed as active in the identity system. Production data did.
  4. Stage your AI evals. Use smaller, cheaper AI models for quick basic checks and routine scenarios, and more powerful models for sensitive cases where failure is more likely.
  5. Build your AI evals into your release process, and adjust your evals based on results. Version your AI evals and do not tune them just to reach a 100% Pass score. As some practitioners have pointed out, optimizing some AI evals for the sake of scoring some queries higher than others can imbalance the system.
  6. Decide on the AI eval set that goes in production and monitor production on a set cadence. Use dashboards, logs, step-by-step records of what the agent did (traces) and do not skip targeted human review.
  7. Use the failing production checks to reinforce your AI eval set. Keep the feedback loop going.

How do you know your AI evals are good enough?

Your AI evals are as proficient as the level of design and sophistication invested before production and after production deployment.

So what does this mean? Let’s look at an analogy first. You need to sit an important test. You go through the handbook and practice tests. You are very accustomed to the test format and you do study diligently. You pass the test with flying colours. Then, you get a job in the field. The job does not have a test format. You need to know how to apply the handbook on the fly. You realize you are not that prepared for the real world. Being a good student and passing tests whose format is familiar is not the same as solving real-life issues.

In the same vein, creating proficient AI evals that are objective enough, smart enough, and flexible enough must be anchored in various good practices which must complement any skilled development team:

  1. Do not confuse passing familiar tests with doing the job well. The fact that a dashboard is green (like in Sam’s case) does not mean that AI evals checked every step of a multi-step task, or that the agent reviews and corrects its own work.
  2. Capture the essence of what good looks like, and what bad looks like end to end and component by component. This means that you do not rely on a single metric, like in Sam’s case, the result, the drop in the number of tickets. You diversify your metrics to catch faulty arguments passed from one component to the other, mis-selection of the appropriate tools by one component, timeouts at various checkpoints, but also measure the entire end-to-end performance over time.
  3. Ground evals in real failures, not synthetic examples and do error analysis. Your agent is designed to solve your specific problems. Many turnkey solutions might be too generic for your needs. An AI eval platform might be useful as infrastructure and management system, but you would still need to nurture it, maintain it, and decide how to run it.
  4. Combine scoring methods. Use fixed rules, LLM-as-a-Judge (a second AI model that grades your agent’s answers), and human review. In other words, decide where you can use calculus (deterministic), AI assessment, exploration and extrapolation (LLM-as-a-Judge), and manual checks (human review). Setting LLM-as-a-Judge requires careful planning and implementation, as this is a project within your project: model selection – data sets – scoring criteria – tests – gated deployment.
  5. Calibrate good results against human interpretation. Irrespective of the scoring method, but especially if you use the LLM-as-a-Judge approach, make sure the judge’s verdicts agree with your human subject-matter experts before you rely on it. Strong AI judges typically agree with human reviewers 80–90% of the time, about as often as two humans agree with each other. How much agreement you need depends on what a missed failure would cost you.

Who is in charge to create evals?

Domain experts and AI teams work in tight collaboration to create the eval datasets. The people who know what a correct result looks like (your domain experts in operations, HR, finance, or compliance) define the test cases and review failures. In Sam’s case, that would have been HR and IT security, together. The good news is that AI evals are a professional skill which can be taught, nurtured and applied successfully.

  1. Agent reliability depends on proficient AI evals. Prioritize ample time for error analysis before choosing an AI eval platform and before choosing whether you want to build or buy. Sam, our agent, was not built to catch the ambiguity and inconsistency between HR offboarding and IT. When HR said an employee had left but the identity system still listed them as active, Sam treated “active in either system” as active. No eval checked for that conflict, so nothing stopped Sam from reactivating the accounts of ex-employees, thus exposing the company to unauthorized and potentially malicious access.
  2. Agent reliability depends on comprehensive AI evals. Group your pre-production agent errors and recalibrate your production failure categories based on ongoing trace monitoring and log alerts.
  3. Agent reliability depends on composite checkpoints. Assess when it is best to use fixed, rule-based checks, which cost almost nothing to run, or AI-model checks, which you pay for per use. Run experiments, observe and improve your AI evals, track scores over time, and version your changes. This composite checkpoint system is what practitioners call an eval harness: a pinned version of the prompt, model, and tools under test; a versioned dataset of cases; a safe test environment so agents can call tools without touching production; the scoring logic itself; and an automatic check that blocks a release when scores drop below your threshold.
  4. Agent reliability depends on trustworthy production data. Integrate your AI evals into your deployment pipeline. Use the production data to provide feedback to AI evals, and use AI evals to monitor production.

Building reliable AI agents requires a heavier initial investment: the right platform, the right AI eval cases, the right labels. But consider the alternative. You ship an AI agent in a matter of weeks and pay a hefty production cost after a couple of weeks of interaction.

What does it cost to skip AI evals?

Without AI evals in place failures will occur. But will the failure be visible or invisible? Will you notice that the system is failing your business when it is too late?

With AI evals in place failures will occur, but most will be caught early, isolated safely, and resolved behind the scenes.

AI evals are a prerequisite for any agentic implementation. Here is what they give you:

  1. You can prove your agents are safe to use. Customers, auditors, and investors increasingly ask for that proof, and your eval results provide it.
  2. Your agents get better over time. Every failure becomes a new test case, backed by real data, meaningful metrics, and combined human and AI scoring.
  3. Fewer quality drops reach production. A release that lowers eval scores is blocked before your customers see it.
  4. Problems in production are caught early and fixed fast, before they turn into incidents, fines, or headlines.

Deloitte refunded the Australian government more than A$97,000 (about US$63,000) of an A$440,000 contract after a report shipped with a fabricated court quote and references to research that does not exist. Alphabet lost about $100 billion in market value in a single day after Bard’s launch demo gave a wrong answer about the James Webb Space Telescope. Next to figures like these, the cost of running evals is a rounding error.

If you are investing in a company that runs AI agents, the same logic applies to due diligence. An agent without evals is an unpriced risk: it can look productive on a dashboard while building up liabilities the numbers do not show. Ask to see eval results over time, not just a demo. Our FAQ below lists the questions to start with.

Before your next agent goes live, or your next investment closes, ask how it is tested. If the answer is a green dashboard, talk to us. Bonsai Labs helps teams design, build, and run AI evals that catch problems before customers, auditors, or investors do.

How do we approach AI evals at Bonsai Labs?

At Bonsai Labs, we start where the problems are, not where the tools are. Before we write a single eval, we study how your workflows fail today: the exceptions, the edge cases, and the contradictions your teams still handle by hand. We build evals from those real failures, not from generic examples, and combine fixed rules, AI judges, and review by your domain experts. Every eval becomes part of your release process, so a change that lowers quality does not reach production. Once the agent is live, we monitor its work and turn every new failure into a new test. The result is an agent whose performance you can see, measure, and explain to your customers, your auditors, and your board.

10 questions executives ask us about AI evals

1. What exactly is an AI eval, in plain terms?

An AI eval is a test that checks whether your AI system does what it is supposed to do. Some evals check hard facts, such as whether the agent picked the right tool or returned a valid record. Others judge quality, such as whether an answer is accurate, complete, and in line with your policies.

2. How are AI evals different from regular software testing?

Traditional software gives the same output for the same input, so a test passes or fails the same way every time. AI models can answer the same request differently, and a change in the model, the instructions, or the data can quietly shift quality. That is why evals run on many realistic cases and keep running after release.

3. What is the difference between offline and online evals?

Offline evals run before release against a fixed set of test cases, and can block a release if scores drop. Online evals run in production as part of monitoring and flag when the agent’s live work starts to slip. You need both: one keeps problems out, the other catches the ones that get through.

4. Who should own AI evals in our organization?

Engineering builds and runs them. Domain experts, the people who know what a correct answer looks like, define the test cases and the pass criteria. Technology and business leaders set the quality thresholds and decide how much risk the organization will accept.

5. How much time and budget should we set aside for evals?

Expect a meaningful share of an agent project to go into understanding failures and building evals, especially at the start. Running evals is comparatively cheap: engineering time to build the test set and harness once, model costs that are usually small per run, and possibly a platform subscription. Compare that with the cost of a single public failure.

6. Can we buy an eval platform instead of building our own?

There are many off-the-shelf platforms that give you the infrastructure to host and run your evals, such as Langfuse, Langsmith, Braintrust.

7. What is LLM-as-a-Judge, and can we trust it?

LLM-as-a-Judge means using a second AI model to grade your agent’s answers against a clear rubric, for qualities fixed rules cannot check, such as accuracy, tone, or whether instructions were followed. You can rely on it once its grades consistently agree with those of your human experts on the same sample. Judges have known biases, such as favoring longer answers, so a human should still review a sample regularly.

8. How often should we re-run our evals?

Run offline evals every time something changes: the prompt, the model, the tools, or the data the agent relies on. Run online evals continuously, or on a set schedule, for as long as the agent is in production. Add every new production failure to your test set so the same problem cannot return unnoticed.

9. How do we know our evals are good enough?

Good evals are built from your real failures, not generic examples, and they check each step of the agent’s work as well as the end result. Their verdicts agree with those of your human experts. If production keeps surprising you, your evals have gaps.

10. What should investors ask a company about how it tests its AI agents?

Ask what the company tests before each release, and whether a drop in eval scores can block a release. Ask how it monitors the agent in production, who reviews failures, and how those failures feed back into the tests. A team that can show its eval results over time understands its AI; a team that can only show a green dashboard may not.

Resources

  1. LLM-as-a-Judge evaluation methods
  2. How to evaluate LLMs and AI agents in production: The Braintrust way
  3. LLM evaluation: methods, metrics, RAG & agent evals guide
  4. Frequently Asked Questions (And Answers) About AI Evals
  5. A Field Guide to Rapidly Improving AI Products
  6. Your AI Product Needs Evals

Let's talk about what AI can do for your business.

Get in touch