See the results
Tidio ranks #1 in G2's AI Agent Evaluations
  • Tell us what you need and we'll design an agent around it
Log in
Tidio
>
Blog
>
AI customer service

Tidio’s Lyro ranks No. 1 among AI CX agents in G2’s 2026 evaluations

Written by: Polina Fomenkova
Updated:
Summarize this post with AI

Key takeaways

  • Highest policy compliance: Lyro scored 92% on policy compliance, 15 percentage points above the average G2 displays.
  • Tied for the top overall score: Lyro passed 40 of 46 tasks for 200 points, level with Retell AI.
  • Above-average accuracy: Lyro scored 87% accuracy, seven points above the displayed average.
  • Actions are verified: G2 checks whether refunds, cancellations, and other actions actually happened in the underlying systems, not just whether the agent said they did.

Every AI agent vendor says their bot resolves tickets. Few let you see what happens when the bot gets a refund request it shouldn’t approve.

G2 decided to check. Its new AI CX Agent Evaluations put nine customer service agents through the same 46 support tasks inside a simulated company, then scored what each agent said and what it actually did.

Lyro, Tidio’s AI agent, came out first in G2’s default “All Metrics” view. It posted the highest policy compliance score of any agent tested.

Use live chat to support your customers in real-time

Learn more about Tidio Live Chat

What are G2’s AI CX Agent Evaluations?

G2’s AI CX Agent Evaluations are hands-on tests of AI customer service agents, run separately from G2’s buyer reviews. Each agent gets the same 46 buyer-informed tasks in a simulated company. G2 then scores accuracy, policy compliance, and whether the agent completed the required actions in real business systems.

Reviews tell you how customers feel about a product. This benchmark tells you how the product behaves under identical, controlled conditions.

That matters because AI agents now do more than answer FAQs. They process refunds, change subscriptions, and update accounts. A wrong answer is annoying. A wrong action costs you money.

How did Lyro score in G2’s AI CX Agent Evaluations?

Lyro passed 40 of 46 tasks and earned 200 points, tying for the highest overall score among nine agents. It scored 92% on policy compliance, the highest of any agent tested, and 87% on accuracy. Both figures sit above the averages G2 displays (October 2026 data).

MetricLyroG2 displayed averageDifference
Tasks passed40 of 46n/an/a
Overall points200 (tied 1st)n/an/a
Policy compliance92% (1st of 9)77%+15 pp
Accuracy87%80%+7 pp

Overall points count passed tasks. Compliance and accuracy are reported as separate measures, so each number tells you something different about the agent.

The tie is worth naming. Retell AI also scored 200 points. G2’s default view lists Lyro first, and Lyro edges ahead on both separate measures: 92% vs. 90% compliance and 87% vs. 86% accuracy.

Why does policy compliance matter for AI customer service?

Policy compliance measures whether an AI agent follows your written support rules, not just whether it completes a request. An agent handling a refund has to know how to process it and whether the customer qualifies. High compliance means fewer refunds you didn’t owe, fewer exceptions you didn’t approve, and fewer angry follow-ups.

Say a customer asks for a refund 45 days after purchase when your window is 30 days. An agent that’s good at actions but weak on policy processes the refund. An agent that’s good at both explains the policy, offers what you do allow, and hands off to a human if the case needs judgment.

That second behavior is what G2’s compliance score rewards. It’s also the score that decides whether you can trust an AI agent with anything beyond FAQs.

Lyro’s 92% compliance reflects how consistently it stuck to the written support policy across G2’s test scenarios.

How does G2 test AI customer service agents?

G2 tests every agent on the same 46 buyer-informed tasks inside a simulated company. The company has customer and order data, a written support policy, and 38 working business tools. An AI judge scores accuracy and policy compliance. Automated checks confirm whether required actions actually happened in the underlying systems.

The last part is the one to pay attention to. Plenty of agents will tell a customer “Your subscription has been canceled.” G2 checks the system to see whether it was.

Here’s the setup at a glance:

  • Tasks: 46, identical for every agent, shaped by what buyers ask for
  • Environment: a simulated company with customer records, orders, and a written policy
  • Tools: 38 working business tools the agent can call
  • Scoring: an AI judge for accuracy and compliance, plus automated checks on actions
  • Independence: scores are separate from G2 buyer reviews

You can read the full methodology on G2’s AI CX Agent Evaluations page.

How did the other AI agents rank?

Nine agents have been evaluated so far. Lyro and Retell AI share the top overall score of 200 points, with LiveAgent close behind at 195. Policy compliance spreads widely, from Lyro’s 92% down to 56%, which shows how differently agents handle the same rules.

RankAgentOverall pointsAccuracyPolicy compliance
1Tidio (Lyro)20087%92%
2Retell AI20086%90%
3LiveAgent19587%86%
4Chipp18582%86%
5BoldDesk18088%79%
6Zoho Desk14564%66%
7Jotform AI Agents14082%56%
8Freshchat10572%70%
9HubSpot Agent Hub9569%67%

What did Tidio learn from the G2 evaluation?

Tidio joined the evaluation at G2’s invitation. The team opened test projects and worked with G2 to explore Lyro’s action limits and find where multi-step workflows broke down. Those failures gave the team concrete fixes to work on, which is the point of testing against someone else’s scenarios instead of your own.

Lyro’s score comes from a process that runs well beyond one benchmark.

We focus on making it easier for customers to build and maintain the knowledge Lyro needs to handle support requests reliably. We also continuously test its performance across hundreds of scenarios, using our own LLM judge to identify weaknesses and guide improvements.

Szymon Piechaczek

Senior ML Product Engineer at Tidio

The work continues after a customer goes live, too.

Our work doesn’t stop at deployment. We help customers understand recurring issues through conversation monitoring and automated analysis of customer satisfaction scores. Working closely with them allows us to turn those insights into product improvements, while our engineers monitor Lyro’s technical performance to identify where it needs to improve.

Kuba Wyszomierski

Senior ML Analyst at Tidio

What’s next for Lyro?

Tidio is expanding Lyro in four directions: proactive shopping assistance for e-commerce, new channels such as voice and SMS, deeper Shopify integration, and conversational setup inside the Tidio dashboard. The goal is an agent that sells as well as it supports, and that takes less effort to configure.

We’re investing in Lyro across the customer journey, with proactive shopping assistance to help e-commerce businesses generate revenue, new channels such as voice and SMS, and deeper Shopify integration. Making Lyro easier to set up is just as important, and conversational setup within the Tidio dashboard is another key area of investment for us.

Sebastian Wilk

Senior Product Manager at Tidio

Lyro already works with the helpdesk and live chat in Tidio, and hands conversations to your human team when a case needs one. In the first half of 2026, Lyro achieved a 72% resolution rate, based on testing conducted by Tidio customers.

See how Lyro handles your policies by testing it on your real support questions

Learn more about AI agents

FAQ

Which AI agent ranks first in G2’s AI CX Agent Evaluations?

Tidio’s Lyro ranks first in the default “All Metrics” view. It passed 40 of 46 tasks for 200 points, tied with Retell AI, and posted the highest policy compliance score (92%) of the nine agents evaluated.

What do G2’s AI CX Agent Evaluations measure?

They measure how AI customer service agents perform on 46 identical tasks in a simulated company. G2 scores accuracy and policy compliance with an AI judge and uses automated checks to confirm that required actions happened in the underlying business systems.

Are G2 agent evaluations based on user reviews?

No. The evaluation scores are separate from G2’s buyer reviews. Every agent is tested in the same controlled environment with the same tasks, data, policy, and 38 business tools.

What is Tidio?

Tidio is a customer service platform for small and mid-sized businesses. It brings live chat, a helpdesk, chatbots, and its AI agent, Lyro, into one place, so your team can handle conversations from your website and other channels without switching tools.

What is Lyro?

Lyro is Tidio’s AI agent for customer service. It answers customer questions using the knowledge you give it, takes actions such as the ones G2 tested, follows your written support policy, and hands the conversation to your human team when a case needs one.

Use live chat to support your customers in real-time

Learn more about Tidio Live Chat

Polina Fomenkova
Polina Fomenkova

Polina is an AI Content Strategist at Tidio with over a decade of experience in tech, SaaS, and product-led growth. She creates research-driven, practical content that helps businesses improve customer communication, scale support with AI, and turn content into a real acquisition channel.