
When we launched YourGPT in 2023, our AI agents were resolving around 60% of the customer requests they were set up to handle.
At that stage, getting AI to answer was not the hard part. Getting it to answer reliably was. Hallucination was still a real concern, especially when the information was incomplete or the customer asked something outside the expected wording. Much of our early work went into grounding YourGPT in business knowledge, improving retrieval, and making the agent recognize when it did not have enough information rather than filling in the gaps.
The technology around us was improving too. Models such as GPT-4 improved reasoning and instruction following, we kept improving the product around them. Better grounding, retrieval, knowledge management, and controls made the answers more dependable for both businesses and their customers. That combination helped push resolution well beyond where we started.
Last year, YourGPT deployments were already reaching around 80% AI resolution in production.
For us, that was an important point in the product’s evolution. AI agents were handling much more than common support questions. They were working across subscriptions, billing, account management, customer verification, and other workflows where resolving the request required access to real business systems.
Some customers began sharing those results publicly. SKNANB reports an 85% customer query resolution rate, alongside faster response times and higher customer satisfaction. Leya AI handles more than 1,000 customer conversations each month and reports 80% of support resolved by AI, including Stripe subscription workflows. Talkmore uses YourGPT for first-line telecom support across subscription and billing enquiries, improving response speed and consistency while allowing human agents to focus on more complex cases.
Reaching 80% also gave us a much clearer picture of the requests that were still unresolved.
The remaining 20% contained more operational complexity. Some requests depended on conflicting knowledge. Others required actions across several systems, stronger customer verification, a successful write to an external service, an approval from a teammate, or better handling of exceptions.
That became a major focus across our product, engineering, AI, and customer teams.
Today, maintained YourGPT deployments can achieve 90%+ AI resolution on eligible repeated requests.
The difference between 80% and 90% is more meaningful than a ten-point increase suggests. At 80% resolution, 20 out of every 100 eligible requests remain unresolved. At 90%, only 10 remain.
So the work between those two milestones effectively addressed half of the unresolved share that remained at 80%.
Customers can now get more support requests handled from start to finish without having to repeat themselves, move between channels, or wait for someone else to take over. That makes support easier and less frustrating because they can get to the outcome they need with fewer steps. Human agents are still there for cases that need judgment or a more personal response, while the AI takes care of requests it can complete reliably. This also gives support teams more time to focus on the cases that actually need their attention.
This update explains what changed, how we measure resolution, where the biggest improvements came from, and how we are now applying what we learned with organizations that want to raise their own AI resolution rates.
YourGPT has progressed from production deployments reaching around 80% AI resolution to 90%+ on eligible repeated requests in maintained deployments. The improvement came from sustained work across five areas: We use a strict definition of resolution. A request counts only when the customer’s intended outcome is completed, required actions succeed and are confirmed by the system responsible for them, AI retains ownership of the conversation, and the same issue does not return within recontact window. For organizations that want our team closely involved in this work, the YourGPT Managed Resolution Program applies the same approach to their own support queue, knowledge, integrations, workflows, and unresolved intents.
The early stages of customer support automation tend to expose broad opportunities.
A large percentage of support volume consists of recurring requests. Once an agent has reliable access to policies, documentation, product information, FAQs, and customer context, some of that volume becomes addressable.
YourGPT agents can work with different training sources, including websites, documents, structured text, and other connected business knowledge. As those sources become more complete, straightforward informational requests become easier to resolve consistently.
The unresolved queue changes as the resolution rate improves.
The remaining requests increasingly depend on operational details:
We began reviewing unresolved requests as a sequence rather than as a single conversational outcome.
For every major request type, we ask:
That gave us a much more useful way to improve the product.
A low resolution rate could now be traced to a knowledge problem, missing action, failed integration, approval rule, verification requirement, or a genuine boundary where the request should remain with a person.

Resolution percentages are difficult to compare when the underlying definitions differ.
A generated response, contained conversation, closed session, and completed customer request are not necessarily the same outcome.
At YourGPT, AI Resolution means the successful completion of an eligible customer request without a human support agent taking full ownership of the conversation.
Five conditions have to be met.
The request must fall within the scope the agent has been configured and connected to handle.
That scope is established before measurement.
Spam, duplicate conversations, greeting-only sessions, and requests outside the approved scope stay outside both the numerator and denominator.
This prevents the denominator from changing simply because certain requests proved difficult.
For an informational question, an accurate answer may complete the request.
For an operational request, the requested operation must also happen.
If someone asks to change an account, cancel a subscription, update an order, create a return, or modify another system, explaining the procedure is only part of the work.
When completion depends on another system, the result from that system matters.
If the agent calls a billing platform, ecommerce system, CRM, internal API, or support platform, the action must return a successful result.
Attempted execution does not qualify as completed execution.
A person can contribute a specific decision or approval while the AI continues managing the request.
A complete human takeover is an escalation, so we do not count it as an AI resolution.
We use a re-contact window.
If the same customer returns with the same unresolved issue during that period, the original interaction does not remain counted as a successful resolution.
The formula is straightforward:
AI Resolution = Resolved eligible requests ÷ Total eligible requests × 100
The important part is applying the same qualification rules consistently.
Several adjacent metrics remain useful. We simply do not use them as substitutes for resolution.
A correct response can complete an informational request.
For an operational request, completion also depends on whatever needs to happen outside the conversation.
Containment measures whether the interaction remained with automation.
It is useful for understanding support load. It does not independently establish that the customer’s intended outcome was completed.
YourGPT can close inactive sessions for queue management.
An inactivity timeout does not contribute to the AI resolution figure.
If an action returns an error or fails to complete, the request remains unresolved.
Keeping these outcomes separate makes the metric harder to improve, but it also makes the result more useful to us.
Five areas had the greatest impact.
As an organization’s knowledge base grows, adding more information is no longer the only challenge.
Keeping that information consistent becomes equally important.
A company may have:
Those sources can gradually drift apart.
A current returns page might say 30 days while an older FAQ still says 45. A temporary campaign document might contain another rule.
The agent needs a dependable source of truth.
YourGPT’s Training Data Anomaly Detection scans training sources for conflicting information and surfaces the affected content for review.
Teams can compare conflicting sources and correct or remove outdated material before it continues influencing customer conversations. This is also why our training best practices emphasize source quality and consistency rather than simply adding more content.
This gave us a practical way to treat knowledge consistency as an ongoing operational task instead of a one-time setup step.
When a specific response needs investigation, Training Audit in Debug Lab exposes the retrieval path.
Teams can inspect:
The underlying source can then be corrected or removed. Training Audit makes that investigation part of the normal debugging workflow rather than leaving teams to infer what the agent retrieved.
That makes debugging considerably more precise.
Real customer conversations also reveal information that was missing from the original knowledge base.
Smart Learning tracks unresolved queries and uses previous conversations to generate suggested knowledge improvements for human review.
The review step is important. Production conversations show teams where coverage is weak, while people remain responsible for deciding what should become trusted business knowledge.
The resulting cycle is simple:
At higher resolution levels, this feedback loop becomes increasingly valuable because the obvious knowledge gaps have already been addressed.
A substantial part of customer service involves an action rather than an answer.
Consider the difference between:
“How can I cancel my subscription?”
and:
“Cancel my subscription.”
The first can often be resolved through knowledge.
The second depends on the billing system.
The same distinction appears across:
Increasing resolution meant giving the agent safe access to more of the systems where those outcomes actually happen.
AI Copilot gives an agent capabilities it can use based on the customer’s request and the context available in the conversation.
It can work with actions and connected functions while deciding how to move a request forward. This is particularly useful for AI support agents that need enough autonomy to handle variations in how customers describe the same underlying task.
Teams define what the agent is allowed to do, while the agent determines how those capabilities should be used.
A company website shows that clearly when visitors type a question instead of picking one from a list. Yotta is a company that runs large data centers, cloud services, and AI infrastructure for other businesses. Their website has to explain that work and get a visitor to the right team. The agent on that site is Ved.
Purusharthvardhan Rathore, Product Marketing Specialist at Yotta Data Services:
YourGPT helps us answer more of the questions visitors ask on our website and generate better-qualified leads.
Our previous chatbot could only handle questions we had already defined. With Ved, visitors can ask questions in their own words, find the right Yotta service, and get the information they need in the same conversation.
Some business processes need more explicit execution logic.
Financial changes, identity-sensitive operations, approval thresholds, ordered workflows, and exception-heavy processes may require teams to define exactly what happens at each step.
AI Studio gives teams control over:
This gives organizations two complementary ways to automate.
AI Copilot can handle requests where autonomous execution is appropriate, while AI Studio workflows give teams more control over processes where the execution path needs to be explicitly defined.
The important part is having the right level of control for the request.
Giving an agent access to actions creates another requirement: the workflow needs to know whether the action succeeded.
For any operation that changes business state, there is an important difference between:
the action was requested
and
the requested change was confirmed.
YourGPT workflows can route actions through success and error paths.
When an action succeeds, the workflow can continue with the confirmed result.
When it does not, the agent can:
This prevents the conversation state from becoming the source of truth for an operation that happened somewhere else.
The system responsible for the change confirms the outcome.
YourGPT agents now execute more than two million operational actions each quarter across API calls, record updates, workflow triggers, ticket creation, and other connected processes.
We deliberately keep that figure separate from resolution.
One customer request can execute several actions, and unsuccessful actions still contribute to action volume.
Action volume tells us how much operational work agents perform. Resolution tells us how much customer work reaches completion.
That distinction became increasingly important as agents started doing more inside business systems.
One of the most important areas beyond 80% involved requests where the AI could perform most of the work but still needed a person to make one decision.
A company may require human approval for:
The agent may already have collected the information, verified the customer, checked the policy, and prepared the next action.
The remaining requirement is human judgment.
Treating every one of these cases as a full support handoff creates unnecessary escalation.
YourGPT therefore provides several levels of human involvement.
Check with Team allows the agent to request a specific answer or decision while continuing to own the customer conversation.
The visitor remains with the agent while the teammate responds behind the scenes. Team notifications can also be sent through Slack or Discord.
This workflow is part of YourGPT Autonomous functions, where teams can combine system functions, APIs, custom logic, and human assistance without forcing every exception into a full handoff.
Assign Member routes the conversation to a teammate while allowing the AI to continue responding.
This is useful when someone needs visibility or responsibility without stopping the active workflow.
Escalate to Human pauses the AI and transfers full ownership to the support team.
That remains the appropriate path when the request itself genuinely requires a person.
Human judgment can be part of an AI-owned workflow without turning every decision into a full escalation. And this helps AI to learn over time from human experience.
That distinction became increasingly important as we worked on the requests remaining beyond 80%.
An overall resolution rate tells us whether a deployment is moving in the right direction.
It does not always tell us what to work on next.
A deployment could resolve one common intent at 97% while another remains at 55%. A blended number can hide the difference.
So we increasingly review resolution by intent.
| Area | What we look for |
|---|---|
| Knowledge | Does the agent have a current, authoritative source? |
| Actions | Can it perform every step required for completion? |
| Verification | Can the workflow confirm that each important action succeeded? |
| Human input | Does it need approval, monitoring, or complete human takeover? |
| Re-contact | Does the customer return with the same issue? |
The result is a much more actionable improvement queue.
A weak intent may need a better source.
Another may be missing an API connection.
A workflow may have inadequate failure handling.
An approval policy may be producing unnecessary escalation.
Some requests may simply belong with a person.
This intent-level view also aligns with how we think about customer support automation: the useful question is not how much support can be automated in theory, but which repeated customer requests can be completed reliably in practice.
Some customer requests span several distinct responsibilities.
A complex workflow may require:
For these requests, specialist agents can separate responsibilities so individual parts of the workflow are easier to control and diagnose.
We do not treat multi-agent architecture as a requirement for every deployment.
A focused request such as an order-status lookup may work better with one well-configured agent.
Specialist agents become useful when the separation creates a practical benefit, such as clearer permissions, isolated failure points, independent verification, or more controlled execution. That is also where agentic AI in customer experience becomes more than a conversational layer and starts coordinating distinct pieces of operational work.
A customer experiences one conversation.
The work behind that conversation can involve several systems.
Take a duplicate billing charge as an example.

A complete workflow may need to:
For a policy question, the workflow can be much shorter.
For billing, commerce, account management, or operational support, more of these stages may be required.
This is why we measure the customer request rather than simply the chat session.
These improvements are reflected in the customer experience. A higher resolution rate means more customers can get what they need in the conversation. For us, that is what resolution should achieve. It should make support easier, faster, and more reliable for the customer, not just keep more conversations with AI.
High resolution should not be interpreted as a goal to automate every customer request.
Some requests should remain human-owned.
That can include:
The YourGPT 90%+ figure applies to eligible repeated requests in maintained deployments.
The achievable rate for any organization depends on several factors.
Structured, repeated requests with clear completion criteria generally provide more opportunities for AI ownership.
The agent needs current and consistent information.
If an outcome lives in another system, the agent needs access to the required workflow or action.
The agent needs enough authority to complete the work the organization has assigned to it.
Important actions need an explicit success or failure signal.
The organization needs to decide when AI can proceed independently, when targeted approval is enough, and when complete human takeover is appropriate.
This is also why our approach to evaluating customer support AI agents focuses on the actual requests and systems inside the support operation rather than a single platform-wide automation number.

AI resolution does not look exactly the same across every industry. The strongest results tend to come from environments where customer requests are repeated, clearly defined, and connected to the systems the agent needs to complete the work.
Across maintained YourGPT deployments, we are seeing 90%+ AI resolution on eligible repeated requests in SaaS, ecommerce, retail, gyms and fitness, and clinics. In banking, telecom, and more complex enterprise environments, some deployments are currently closer to 85%.
| Industry | AI resolution | Handled |
|---|---|---|
| SaaS | 90%+ | Account support, subscriptions, billing, cancellations, product questions, Product Learning |
| Ecommerce | 90%+ | Order status, returns, refunds, shipping, product questions |
| Retail | 90%+ | Product information, availability, orders, returns, store policies |
| Gyms & fitness | 90%+ | Memberships, bookings, plans, cancellations, routine enquiries |
| Clinics | 90%+ | Appointments, scheduling, service information, patient enquiries |
| Banking & financial services | ~85% | Account enquiries, product information, service requests, policy-driven support |
| Telecom | ~85% | Billing, subscriptions, plans, service enquiries |
| Complex enterprise support | ~82% | Multi-system requests, approvals, permissions, and exception-heavy workflows |
The difference often comes down to how much of the request the agent can complete on its own. In e-commerce, for example, an agent can handle a large share of repeated requests when order data, shipping information, returns, and refund actions are connected. In clinics or gyms, appointment and membership workflows are similarly structured and repeat frequently.
Banking, telecom, and complex enterprise support often involve more permissions, approvals, policy checks, or systems that have to work together. Those deployments can still resolve a large share of customer requests, but more cases may require additional verification or human input.
The independent Resolution Index report places YourGPT in the 80% to 90% AI resolution range, which provides useful external context alongside the performance we see across different deployment types.
The important point is that there is no single resolution rate for every YourGPT deployment. The achievable level depends on the requests being handled, the systems connected to the agent, the actions it is allowed to perform, and where the organization chooses to keep people involved.
Moving beyond 80% changed how we think about AI resolution. The remaining requests are usually less about answering a question and more about whether the agent has the right knowledge, access, permissions, and support to complete the work.
Connecting documents and websites is a starting point. Policies change, old content remains indexed, and customers raise questions that were never covered in the original knowledge base.
Keeping resolution high requires ongoing maintenance so the agent works from current, consistent information.
Many requests can only be completed if the agent can access another system.
An agent may know how a refund, cancellation, order update, or account change should work, but resolving the request depends on being able to perform that action in the relevant billing, commerce, CRM, or internal system.
For operational requests, attempting an action is not enough.
The agent needs to know whether the change actually succeeded before it tells the customer the request is complete. That confirmation should come from the system where the action happened.
Some requests need one human decision, while others genuinely need a person to take over the conversation.
Treating those cases differently helps the AI continue handling work it can complete while bringing in a teammate only for the judgment, approval, or exception that requires one.
The unresolved requests become more useful when resolution is already high. With fewer failures to review, it is easier to see what is actually stopping the agent from completing the job.
Sometimes the problem is simple. The agent may be missing the right information, lack access to an action, run into a permission issue, or depend on an integration that does not behave reliably.
Resolved conversations can still show where the experience needs work. A request might reach the right outcome but take more steps than necessary, pause for an approval, or ask the customer for information the business already has.
That gives teams a clearer picture of what to improve next, not only where conversations fail, but where successful ones can be made faster and easier.
The exact ceiling will vary by organization, but the process is repeatable.
Review real support history and identify the requests your team handles most frequently.
Focus first on work that has a defined outcome.
Do this before measuring.
For example:
This gives every workflow a clear completion criterion.
For every important intent, define where the agent should get the information it needs to complete the request.
A return request may depend on the latest return policy, while an order issue may require both live Shopify data and current shipping rules. When the same information exists across different pages, documents, or systems, the agent should have a clear source to rely on.
This becomes especially important as the knowledge base grows. YourGPT’s AI agent best practices cover how to maintain reliable training data, test responses, and keep the agent’s knowledge accurate over time.
For every operational intent, map what has to happen from the customer’s request to the final confirmed outcome.
A refund may require customer verification, an order lookup, a policy check, a payment action, and confirmation from the billing system. An order change may involve Shopify, shipping data, and an internal approval.
This makes it easier to see exactly where a request can stop. If the agent is missing one required action, permission, or system connection, that gap directly limits how much of the intent it can resolve end to end.
For every important action, specify:
Do not make every exception a full escalation.
Decide which cases require:
Use actual historical conversations rather than generic test messages.
Include normal requests, incomplete information, verification failures, action errors, approval cases, and edge conditions.
The blended number gives you direction.
Intent-level performance tells you what to fix.
Prioritize the request types that contribute the most unresolved volume.
Unresolved conversations should feed the next iteration.
A recurring unresolved request may point to:
This is how resolution becomes an ongoing operating discipline rather than a launch metric.
The work behind a 90%+ resolution rate is broader than configuring an AI agent.
It requires understanding the support operation, defining the right scope, maintaining business knowledge, connecting operational systems, designing workflows, setting approval boundaries, testing against real requests, and continually reviewing what remains unresolved.
For organizations that want our team and our trusted partners to be directly involved, we created the YourGPT Managed Resolution Program.
The objective is clear:
Work closely with your organization to build toward 90%+ AI resolution across eligible repeated requests that are suitable for AI ownership.
We begin with the support operation itself.
Our team reviews historical conversations to establish:
We work with your team to decide:
We review the sources behind important customer requests and identify:
For operational requests, we map the work required for completion and connect the appropriate systems through YourGPT’s autonomous actions, Functions, AI Studio, APIs, and existing integrations.
The same Functions framework used for APIs, custom logic, and Check with Team becomes part of that resolution architecture where required.
Our team works with your organization on:
We test against real request patterns, including the cases most likely to limit resolution:
After deployment, we review the intents still generating unresolved volume.
Those requests determine the next set of improvements.
One organization may need stronger knowledge coverage.
Another may need billing, commerce, or CRM actions.
Another may have excellent workflow coverage but too many full handoffs caused by approval rules.
The program is designed to find those constraints and work through them with your team.
Organizations interested in this level of hands-on resolution work can send their deployment details through that form.
Last year, reaching around 80% showed us that AI agents could take meaningful ownership of customer support.
Getting beyond 90% required going deeper into the parts of the request that happen outside the final response.
The knowledge has to be dependable.
The required action has to be available.
Execution has to succeed.
The system responsible for the outcome has to confirm it.
A person needs to enter at the right point when judgment is required.
And unresolved requests need to tell us what to improve next.
Those principles now influence how we build YourGPT across knowledge, actions, AI Studio, autonomous agents, human collaboration, self-learning, and analytics.
They also reflect the broader direction of the YourGPT AI agent platform, where customer-facing agents can combine business knowledge with real operational actions across support, sales, and business workflows.
As Rohit Joshi, Co-founder of YourGPT, puts it:
Resolution isn’t a chat metric. It’s an operations metric. The goal isn’t to have more conversations with AI. It is to help AI resolve repetitive queries and complete complex customer requests, either autonomously or with targeted human assistance. Customers care about whether their request actually gets handled.
That remains the direction we are working toward.
The 90%+ milestone matters because more eligible customer requests are reaching completion.
The remaining unresolved requests matter just as much because they show us where to work next.
When a customer asks for something the AI is designed to handle, the aim is to complete the request, not just answer it. In properly configured YourGPT deployments, AI can resolve more than 90% of eligible repeat requests without a human taking over. We count a request as resolved only when the customer gets the outcome they need and any required action is completed successfully.
A request counts as resolved when it falls within the eligible scope, the customer’s intended outcome is completed, required actions succeed, and AI retains ownership of the conversation.
Yes, when the customer only needs information and the response completes the request. For operational requests such as refunds, cancellations, or account changes, the required action must also be completed successfully.
Cost does not scale directly with capability. Adding more actions, verification steps, and system connections does not produce a proportional increase in spend.
The cost of resolving a request depends on how complex that request is. A policy question and a multi-system billing dispute involve different amounts of work, and pricing that treats them the same does not reflect the actual cost of running a support operation.
Most providers in this category charge separately for resolution, adding an outcome-based fee on top of subscription or seat-based pricing. YourGPT does not do this. Resolution is not billed as a separate line item. The cost of a deployment does not increase simply because more requests are completed instead of escalated.
This matters as resolution improves. A model that charges per resolved request means an organization pays more as its support automation becomes more effective. That works against the purpose of automation, which is to reduce the cost of the work, not move where the cost is applied.
The more relevant question is not the cost of a single resolution, but the cost of the deployment overall, and how that changes as more requests are completed rather than escalated.
Yes. A teammate can provide an approval or specific input while the AI continues handling the conversation. A complete human takeover is treated as an escalation.
The achievable rate depends on the types of requests being handled, knowledge quality, available actions and integrations, permissions, approval requirements, and which requests the organization chooses to keep human-owned.
Share your queue. We map what still blocks completion.

TL;DR Claude Fable 5 launched on June 9, 2026, was pulled offline three days later under a US export control order, and returned on July 1 after the order was lifted. For support teams, Fable 5 is best suited for long-horizon tickets across billing, CRM, shipping, and documents not basic FAQ deflection. Its biggest support […]


At YourGPT, we’ve always believed businesses need one platform for customer support, sales, and engagement. These are not separate parts of the customer journey. They shape how businesses communicate, build relationships, and move conversations forward. That belief continues to guide how we evolve YourGPT, building toward a more connected way for teams to manage communication, […]


TL;DR YourGPT Copilot SDK is an open-source SDK for building AI agents that understand your application state, user context, and what is happening inside your product. Instead of working like isolated chat widgets, these agents connect directly with your product and can take real actions within existing workflows. This helps teams build contextual AI experiences […]


Happy New Year! We hope 2026 brings you closer to everything you’re working toward. Throughout 2025, you’ve seen the platform evolve. We shipped the AI Copilot Builder so your AI could execute actions on both frontend and backend, not just answer questions. We added AI assistance inside Studio to help you generate workflows without starting […]


Grok 4 is xAI’s most advanced large language model, representing a step change from Grok 3. With a 130K+ context window, built-in coding support, and multimodal capabilities, Grok 4 is designed for users who demand both reasoning and performance. If you’re wondering what Grok 4 offers, how it differs from previous versions, and how you […]


OpenAI officially launched GPT-5 on August 7, 2025 during a livestream event, marking one of the most significant AI releases since GPT-4. This unified system combines advanced reasoning capabilities with multimodal processing and introduces a companion family of open-weight models called GPT-OSS. If you are evaluating GPT-5 for your business, comparing it to GPT-4.1, or […]
