AI Reality Case Studies

Almost Magic Tech Lab seal: Clarity over cynicism. Less noise, more nerve.

The complaint the queue called closed

An AI Reality case study - Customer Service and Support

Fictional composite | Free reader edition, 1 October 2026

The people, organisation, numbers, messages and internal exhibits below are invented for teaching. They are not AMTL client evidence. Real cases cited later illustrate distinct accountability and measurement issues, not Rae's incident.

Monday: a good-looking queue

The Monday dashboard gave Imani good news. The support team had introduced an AI assistant that sorted messages and drafted replies for human agents to check. The backlog was falling. Most cases now showed a closure within two days, and managers wanted the pilot extended to the evening shift before the seasonal spike. Agents liked the faster first draft; the operations manager wanted the queue under control without an extra shift of reviewers.

Imani opened a case marked closed. Rae had written that a delivery charge appeared twice. The AI assistant drafted a courteous explanation of how card holds sometimes appear, and a human agent sent it after a quick read. Rae answered: “Thanks, that helps. I still need the second charge reversed, and the bank says both have settled.” The AI assistant classified “Thanks, that helps” as a positive signal and suggested closure. The human agent accepted the suggestion while clearing a crowded queue. The agent did not read Rae's second sentence carefully enough and did not request a payment review. The status panel said “no pending action”, because the closure rule looked for the selected status, not the customer's last requested action.

Rae called again on Thursday. Her earlier words remained in the record, but the case number said “resolved”. Her bank's statement that both charges settled was not yet North Quay's payment-system finding. The finance team needed to check whether there were two captured charges, a hold and a capture, or a different explanation. The service lead wanted a clearer human review prompt. Operations wanted the evening rollout. Imani could not treat a classifier suggestion as the cause of everything: the human agent clicked close, and the rule that allowed closure without recording the last requested action predated the assistant. AI made it faster and easier to repeat.

Exhibit 1. The case record visible on Thursday

Step Recorded event Who acted What remains open
Monday 9:05 Rae reports duplicate charge Customer Whether both entries were captured payments
Monday 9:16 Explanation of possible card holds drafted AI assistant Whether the explanation fits Rae's account
Monday 9:20 Draft sent after quick read Human agent No payment review opened
Monday 10:02 “Thanks, that helps. I still need the second charge reversed...” Customer Explicit action request remains
Monday 10:04 Positive sentiment; closure suggested AI assistant Suggestion is not a payment finding
Monday 10:06 Case marked closed Human agent Rae's requested action absent from pending queue
Thursday Rae calls again Customer Remedy and accurate status needed

All times and messages are fictional. “The customer was polite” and “the customer's problem was resolved” are different claims.

Exhibit 2. The local closure rule

Current rule Gap Proposed test before closure
An agent may accept the suggested status when the panel shows no pending task The panel does not require the last customer request to be named Write the last requested action, its owner, evidence of completion or why it does not apply, before closing
Closed cases leave ordinary follow-up queue A mistaken closure removes work from view Route any new request after a thank-you to human review; reopen if unresolved
Performance report counts time to closure A faster status can look like a better outcome Track repeat contacts and unresolved requests beside closure speed

The rule's flaw is independent of the tool. Removing the assistant without changing the rule leaves a manual version of the same problem; keeping it while preserving the rule may scale the flaw to another shift.

Exhibit 3. A bounded audit, not a population verdict

Imani can allocate two reviewers for a combined sixteen hours before any revised evening-shift release. At eight minutes per case, theoretical capacity is 120 cases before rework. She proposes 60 cases selected from closures with a “thanks” followed by further text, 30 payment disputes and 30 randomly drawn from other recent closures. These groups need de-duplication where they overlap; the actual distinct count may be under 120. The sample deliberately over-represents suspected failure modes. It can find cases needing action and test whether a revised rule catches them; it cannot yield a simple overall failure rate for all closed cases. A separate random sample would be needed for an estimate of prevalence.

Measure What management saw What Imani still needs
Time to closure Most cases within two days Time until the last requested action was completed
Backlog Falling Work missing from the queue because of premature closure
Rae's repeat contact One visible return Whether similar mixed-sentiment requests recur elsewhere
Refund status Not started Finance check, authorised decision and accurate customer update

The numbers are invented planning assumptions, not survey results or measured rates.

A documented parallel, not Rae's case

In Moffatt v Air Canada (2024), a Canadian tribunal found the airline responsible for misleading information its chatbot gave about a bereavement fare. That was an incorrect answer, not a complaint marked closed. Its link to Rae's case is organisational accountability: an automated message or classification inside a company's service process does not shift the duty to explain and fix the customer outcome to the software. It does not prove a refund is owed to Rae. https://decisions.civilresolutionbc.ca/crt/crtd/en/525448/1/document.do

UK Department for Work and Pensions research reported that some interviewed customers believed they had made a formal complaint when the department had not classified it as such. That is not AI research. It is a parallel for the gap between an internal status and a person's experience. Rae thinks her request is pending; the dashboard says closed. Both sources make the measurement question worth asking, but neither supplies a failure rate for this fictional pilot. https://www.gov.uk/government/publications/research-examining-customers-experiences-of-dwps-complaints-process/research-examining-customers-experiences-of-dwps-complaints-process

The case in numbers

Times and capacity figures come from the case record or follow from it. Three figures are invented teaching assumptions, marked in the last column so you can replace them with your own.

Metrics for this case
MeasureFigureWhy it mattersBasis
Rae's report to closure61 minutes (9:05 to 10:06)The case closed about an hour after it openedDerived from Exhibit 1
Rae's last message to closure4 minutes (10:02 to 10:06)Little time to read her second sentenceDerived from Exhibit 1
Time before Rae called again3 days (Monday to Thursday)A closed case does not stop a customer waitingDerived
Rae's duplicate delivery charge$14.95Small, but the case is about the principleInvented
Median time to closure3.4 days down to 1.6 days (down 53%)The number management sees. It measures closure, not resolutionInvented
Open backlog over four weeks410 down to 260 cases (down 37%)Falling, in part because closed cases leave the queueInvented
Reviewer capacity for the audit16 hours (2 reviewers, 8 hours each)Under half of one person's working weekStated; comparison derived
Cases the audit can cover120 at 8 minutes each, before reworkThe sample may be smaller after de-duplicationStated
Planned audit sample60 + 30 + 30 = 120 casesChosen to find failure modes, so it cannot give a failure rateStated

The invented 53% is a speed figure. No number on this page measures how many customers got what they asked for, and that gap is the case.

Critical evaluation

Question Working answer What remains open
What happened? The AI suggested closure; a human accepted it after an unresolved request. How often recent closures hide an unfulfilled ask.
What does the result support? The queue closed faster and Rae returned. Whether customer outcomes improved, or status handling changed.
What could explain it another way? Rushed human reading, status-rule design and AI focus on positive sentiment all matter. Which control would catch the next mixed-sentiment message.
What test comes next? Audit a targeted sample against the last requested action, then validate a changed rule. Random prevalence estimate and true capacity for expansion.

DECISION MAP
Customer asks -> AI draft -> human sends -> customer replies -> AI suggests status -> human accepts -> case exits queue. Place the missing review at the customer's last action request. The diagram shows the mechanism; the table below shows the choices. They are not the same thing.

Figure 1. The closure handoff: where politeness becomes a status.

Decision handoffs and the missing evidence check in Customer Service and Support

The decision room: Thursday, 2:30 pm

Imani has to tell operations by 3 pm whether to train the evening shift on the current process. Rae also needs a truthful update today. Choose a path and say who checks her payment record, who can authorise a reversal, what the customer is told now and what audit result would reverse your rollout decision.

Choice Immediate gain Immediate cost Second-order risk
A. Leave Rae closed; expand evenings Faster visible queue reduction No remedy for Rae; current rule unchanged More unresolved work disappears from ordinary follow-up
B. Reopen Rae; remove AI assistance Human review for Rae and no further AI closure suggestion Agents lose useful triage and drafts; backlog may grow The old closure rule still permits manual premature closure
C. Reopen Rae; retain drafts but gate closure Rae gets a payment review; useful sorting continues Sixteen review-hours plus delay to evening rollout New rule must be checked, not simply added to the interface

C is not free. If review finds many hidden requests, more staffing and a longer pause may be required. If payment records show only a hold, the team still owes Rae an explanation, not a reversal promise. Decide before reading the follow-up.

OPERATING PRINCIPLE
A case is not resolved because the customer was polite. It is resolved when the remaining request has an answer.

Separate teaching epilogue: open only after choosing a path

Imani chose a bounded C, not because it split the difference. A left a visible customer request unanswered and scaled the same closure rule. B removed the assistant but left the human-acceptance rule untouched. She reopened Rae's case, asked finance to compare payment captures with the bank claim, and gave the authorised refunds team the result. The human agent told Rae that the charge was under review, supplied a reference and a time for the next update, and did not say the bank had already credited anything.

Imani paused the evening rollout. Two reviewers examined the planned targeted sample and recorded each last requested action, whether it had an owner, and whether the case was actually resolved. The service lead altered the closure screen so a human had to record the last customer request and either completion evidence or a reason it did not apply. Mixed-sentiment replies went to review rather than automatic closure advice. Finance, not the AI assistant or frontline agent, made the payment decision. Operations accepted a one-week delay, with a decision meeting rather than a vague “we'll see”.

At thirty days the team found that closure speed had overstated resolution in some reviewed cases; the targeted sample could not establish a population rate. Reopening hidden work raised the unresolved count. That was not necessarily service worsening: some work had become visible again. At ninety days the assistant still helped draft and route, but the report placed closure time next to repeat contact and outstanding requested actions. An agent could override the assistant, but had to leave a reason. A different audit result, or no reviewer capacity, might have justified B and a longer delay. The question remains: what did the customer ask to happen, and where is the evidence that it happened?

Sources and evidence limits

Original fictional composite drafted 29 September 2026. The real tribunal decision and DWP research above are parallels with stated limits. Further context: NIST Generative AI Profile https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence ; ACCC false or misleading claims https://www.accc.gov.au/consumers/advertising-and-promotions/false-or-misleading-claims . No source verifies Rae or the invented audit. Almost Magic Tech Lab Pty Ltd.

Facilitator guide

A private 90-minute session kit: run sheet, stakeholder questions, model analysis, transfer worksheet and quality check.

Request the facilitator guide