A dead battery, a broken umbrella and a plastic package can become administrative questions the moment someone tries to throw them away. The answer may depend on chemistry, dimensions, attached materials, the resident’s district and the day of the collection calendar. The question may also arise at 9:30 at night, when the object is already sitting beside the kitchen bin and the municipal office is closed.

Miyoshi City in northern Hiroshima Prefecture spent July testing whether a voice-based artificial-intelligence system could occupy that gap. From July 1 through 31, residents could telephone an AI service about waste categories, disposal methods, collection days and districts. The system, a product called Voice AI supplied by Verbex Inc., answered first and could offer a transfer to a city employee when needed.

The experiment produced 394 calls and a seemingly irresistible set of percentages: 49.5% outside a broad definition of city opening hours, 71.5% answer accuracy and 74.4% “AI handling completion.” Verbex’s corporate announcement also reported a 2.9% incidence of answers not grounded in fact.

Those figures are accurate only when their denominators and definitions travel with them. Miyoshi’s own 11-page report makes that unusually clear. It says the 74.4% figure must not be read as the share of inquiries solved by AI. It says no user-satisfaction survey was conducted. It says the city lacked the baseline and follow-up data required to calculate staff savings. It says cost-effectiveness could not be evaluated.

394 callsThe full month’s usage count. Total connected time was about 7 hours and 32 minutes.
68.8 secondsThe average call. The median was 61.5 seconds, and 91.4% ended within two minutes.
93 of 130The basis for the reported 71.5% accuracy rate. Sixty-one sampled calls were excluded from this calculation.
5 of 175Assessable calls in which reviewers found information not grounded in fact—a 2.9% incidence.
Do not turn 74.4% into “the AI solved three quarters of the calls.” The numerator included 79 calls completed by AI alone, 10 successful transfers to staff and 39 cases in which no answer or transfer was achieved but the AI’s conduct itself was judged not problematic. The AI-alone completion rate was 41.4% of the 191-call review sample.

A small service with a revealing dataset

The calls totaled 7 hours, 32 minutes and 1 second. Nearly half—48.2%—ended within one minute; 91.4% ended within two. These were largely quick, concrete questions rather than open-ended conversations about public policy.

Using automatically generated call summaries and a rules-based classification method, the city placed 293 calls, or 74.4%, in “item classification and disposal method.” Another 69, or 17.5%, concerned collection dates or districts. Seven involved drop-off or bulky-waste procedures, two concerned unrelated administrative topics, and 23 contained no clear or valid question. The report says batteries and rechargeable batteries, plastics and containers were among the recurring themes.

The classification itself has limits. It assigned one principal category to each call from machine-generated summaries; it was not a human classification of every complete audio recording. That distinction matters because a single conversation may contain several questions.

Demand also fell sharply after the launch. Seventy-two calls arrived on July 1 alone—18.3% of the monthly total—probably reflecting publicity and people trying the new number. Excluding that first day, the daily average was 10.7. Miyoshi explicitly warns that the pattern cannot establish continuing demand. Multiplying 394 by 12 would create a number the evidence does not support.

Still, the trial confirms a real workflow: residents will use a telephone to ask brief, repetitive and highly local questions. Each one may occupy only a minute, but a municipal worker must stop, identify the item and district, locate the rule, explain it and then return to another task. The commercial opportunity lies in that accumulation of interruptions. The public-service test is whether automation can remove them without transferring the risk of a wrong answer to the resident.

After hours, by two different clocks

Miyoshi published two useful—and different—definitions of “after hours.” Its summary webpage treats Monday through Friday from 8:30 a.m. to 5:15 p.m., plus Saturday from 8:30 a.m. to noon, as open hours. On that basis, 195 of 394 calls, or 49.5%, came outside open hours.

The detailed report instead creates a standard “weekday daytime” category: weekdays from 8:30 a.m. to 5:15 p.m., excluding the July 20 public holiday. By that definition, 182 calls came during weekday daytime and 212 came outside it. The resulting after-hours share is 53.8%. The report separately counts 79 calls on Saturdays and Sundays.

There is no need to choose the more dramatic percentage. The boundary changes because Saturday morning and the public holiday are treated differently. Both calculations support a narrower conclusion: a substantial portion of the inquiries occurred at times not covered by an ordinary weekday desk.

The public value is not simply that software can “work longer.” It is that the moment a resident has a question no longer has to coincide with the moment a municipal employee is available.

The denominator test

Miyoshi did not manually score every one of the 394 calls. It divided the month into four periods and selected 50, 50, 50 and 41 calls, producing a 191-call review set. It then removed calls that could not be judged from the denominator of each relevant metric. As a result, percentages printed beside one another do not describe the same universe.

Published metricWhat was countedWhat it does—and does not—mean
97.1% free of unsupported answers170 of 175 assessable calls contained no answer ungrounded in fact. Five did; 16 were not assessable.The corresponding observed incidence was 2.9% among assessable reviewed calls, not among all 394 calls.
71.5% answer accuracy93 correct answers among 130 calls whose correctness could be determined. There were 37 wrong answers or answer failures.It is not 71.5% of the 191-call sample. Forty-eight calls lacked required registered information; 13 more could not be judged because conversation failed or for similar reasons.
74.4% AI handling completion128 of 172 calls with a judgeable handling outcome.This composite includes AI-only completion, successful human transfer and some cases with no answer but no handling fault.
41.4% AI-only completion79 of all 191 reviewed calls ended without transfer to staff.It is the cleanest published measure of calls finished by AI alone, but it is not a satisfaction score and does not certify every answer as correct.

The 61 calls left out of the accuracy calculation are as informative as the 130 that remained. For 48, the information necessary to answer had not been registered in the system. Thirteen more could not be judged because the conversation did not form or for related reasons. This is not merely a problem of making a language model “smarter.” It is a problem of turning municipal rules into governed, current and machine-retrievable knowledge.

The detailed report defines a hallucination as the AI providing content that is not based on fact. Reviewers found five such calls among 175 that could be assessed for that behavior. The public document does not disclose the five exchanges, their subjects or the seriousness of the possible consequence. It also appropriately excludes telephone numbers, call IDs and individual summaries that could identify calls.

Garbage is local law disguised as an everyday object

Waste sorting looks like an attractive starter task for AI because many questions repeat. It is also a difficult one because the correct answer is not just a fact about an object.

A plastic item may be packaging or a product. Dirt, attached metal, material composition or physical size can change its category. Rechargeable batteries require safety instructions that ordinary dry cells may not. A bulky item may need an appointment or direct delivery. A collection date depends on an address and district. An answer that would be acceptable in another Japanese city may be wrong in Miyoshi.

This local variation is structural. Under Japan’s Waste Management and Public Cleansing Act, responsibility for ordinary municipal waste rests fundamentally with municipalities. The national government sets the broad legal framework, but a resident needs the rules, calendar and acceptance conditions of the place where the item will actually be collected.

That makes the system’s job more demanding than fluent conversation. It must recognize local item names and place names, connect them to Miyoshi’s authoritative data, ask for missing attributes, and refuse to improvise when the answer remains uncertain.

A telephone at the end of 126 years of municipal responsibility

The institutional history behind a garbage hotline reaches back well before telephones became ordinary household objects. Japan’s 1900 Filth Cleansing Law—enacted after infectious-disease crises including cholera and plague—established waste and human-waste removal as an administrative service carried out by cities and specified towns and villages.

The Public Cleansing Act of 1954 expanded the implementing responsibility across municipalities and strengthened the postwar sanitation system. Rapid economic growth then transformed both the volume and composition of waste. In the 1970 “Pollution Diet,” lawmakers replaced that framework with the Waste Management and Public Cleansing Act, distinguishing municipal solid waste, for which municipalities remained responsible, from industrial waste governed by producer responsibility.

By 2000 and 2001, the Basic Act on Establishing a Sound Material-Cycle Society and a suite of recycling laws were shifting the objective beyond sanitary disposal toward reduction, reuse and recycling. Every extra category and material stream made accurate household instructions more valuable—and the local rulebook more intricate.

1900 — The Filth Cleansing Law makes waste collection an administrative responsibility for cities and designated municipalities.

1954 — The Public Cleansing Act expands municipal responsibility and the sanitation system.

1970 — The Waste Management and Public Cleansing Act separates municipal and industrial waste responsibilities.

2000–2001 — The sound material-cycle framework and major recycling laws broaden the policy from disposal to resource circulation.

July 2026 — Miyoshi tests a 24-hour voice entrance to its municipal waste knowledge.

Seen in that history, the AI line does not outsource the city’s public responsibility. It creates another doorway into information for which the city remains accountable. The crucial questions are therefore institutional: Who controls the approved answers? How quickly do calendars and categories get updated? Which uncertainty triggers a transfer? Who reviews errors, and how are residents protected from repeated ones?

Why the telephone matters in Miyoshi

Miyoshi covers 778.18 square kilometers. The present municipality was created in 2004 by combining the former Miyoshi with four towns and three villages. A city planning document places the resident-register population at about 47,400 in September 2025 and notes that population decline has been particularly pronounced in the former town and village areas.

That geography gives the experiment significance beyond a technology demonstration. A web search requires a resident to know what phrase to enter, select the correct page and navigate a local calendar. A telephone lets someone begin with the object in ordinary language. It may be particularly useful where distance, age, device preference or digital confidence makes a conventional web interface less convenient.

But the same geography increases the cost of recognition errors. If the system mistakes a district name, it may return the wrong collection day. Miyoshi’s report identifies both item names and district names as areas where speech recognition failed. Accessibility and accuracy are not competing side issues here; they meet in the same spoken place name.

Four places where a voice answer can fail

A voice-AI call contains several technical stages. Speech recognition converts sound to text. Language processing interprets the intent and details. A retrieval or response system consults registered municipal information and composes the answer. Text-to-speech turns that answer back into sound.

Miyoshi found problems across this chain. The system sometimes failed to recognize an item or district. Some questions concerned information not yet registered. A resident’s brief acknowledgment could be mistaken for an interruption, cutting off the AI’s explanation. Some waste categories were incorrectly stated. Transfers needed better rules and follow-through.

Verbex describes its commercial platform as combining proprietary speech models with outside large language models selected for each use. Its company history says the project began with a basic IVR prototype in 2017, added a Rasa-based dialogue engine in 2022, introduced a generative-AI engine in 2024 and began broader business deployment in Japan and other countries in 2025. Those are company statements, not findings established by Miyoshi’s evaluation; the city report does not disclose the precise model stack used in each call.

The trial therefore tested something more exacting than whether synthetic speech can sound natural. It tested whether a modern voice system can turn imperfect, colloquial and locally specific speech into an answer safe enough to act upon.

The ability to say “I don’t know”

In a social conversation, a plausible mistake may be awkward. In public administration, it can change someone’s conduct. A wrong battery instruction can contribute to a fire. A wrong collection day can leave waste at a neighborhood station. A wrong drop-off instruction can waste a resident’s trip and time.

Miyoshi’s report sets out a standard that deserves wider attention. In Japan.co.jp’s translation, it says: “Quality improvement means not only widening the range of questions the AI can answer, but also refraining from forcing an answer when confidence is low and directing the matter to staff confirmation.”

This reverses a common product incentive. A system optimized for completion will try to produce an answer. A system optimized for safe public service will sometimes ask again, identify the authoritative basis, limit its claim or stop and transfer.

For an administrative AI, “I cannot answer that reliably” is not necessarily a failed interaction. It can be the most accurate and useful answer available.

The five unsupported responses should therefore be treated as more than a blemish on an otherwise high percentage. They are the cases from which designers can determine when the system should not speak. Miyoshi lists the next priorities: improve dictionaries for items and districts, expand registered information by frequency and risk, strengthen verification against authoritative answers, and route difficult questions to people.

No, the trial did not prove a 95% reduction in work

Only 20 of the full 394 calls ended in a staff transfer, or 5.1%. That creates a tempting sales claim: the AI kept 94.9% away from workers. Miyoshi expressly rejects that inference. Its summary says this result alone cannot establish that employee workload fell by 95%.

The city had not measured the pre-trial time spent on comparable calls. The report did not include employee time after a transfer, time spent reviewing transcripts and errors, or the labor required to build and update the answer base. A system label saying the user ended a call—187 cases—could not distinguish successful completion from abandonment. The 184 calls ended by AI were likewise not automatically labeled as solved.

There was no post-call satisfaction survey, so a 60-second call cannot automatically be called convenient. The resident may have received the answer quickly, or may have given up. The city also lacked the operating-cost and time-savings data needed for a return-on-investment calculation.

What the trial demonstrated is narrower and still worthwhile: residents used a 24-hour voice channel; some inquiries were completed without staff transfer; repeated topics could be identified; and specific recognition, knowledge and conversation-control failures could be measured. What it did not demonstrate was a cheaper service, a happier caller or a permanent staffing reduction.

What a second trial should answer

Japan’s 2026 priority plan for a digital society places pressure on local-government resources beside the rapid development of generative AI and AI agents. That national policy explains why experiments like Miyoshi’s will multiply. It does not validate any particular vendor or deployment. Each service still needs evidence tied to the public task.

Five measurements for the next phase
  1. A pre-AI baseline: comparable call volume, average employee handling time and after-call work.
  2. A resident outcome: whether the answer helped, whether the disposal action was completed and whether the resident called again.
  3. Risk-weighted accuracy: not only how many answers were wrong, but whether each error concerned an inconvenient calendar detail or a safety-critical battery instruction.
  4. Total staff work: transfers, quality review, knowledge-base maintenance and correction of bad guidance.
  5. Cost-effectiveness: communications, licensing, integration and oversight costs compared with time saved or reassigned.

A stronger evaluation would also specify how calls are sampled, publish anonymized categories for unsupported answers, test whether performance differs by age, speech pattern or district name, and report what happened after a transfer. Accuracy should be compared using the same method before and after tuning.

Those measurements might make the service look better. They might also show that some question classes are too rare, too dangerous or too complex to automate. Either result would be useful. A pilot is supposed to reduce uncertainty, not merely produce a launch statistic.

The most credible outcome was the admission of limits

Trash sorting is not a glamorous frontier for artificial intelligence. That is why it is a revealing one. The questions are ordinary, the answers are locally governed, and mistakes must be handled in the physical world. Residents gain something concrete if they can ask at night. Workers gain something concrete if repetitive interruptions fall. The city gains something concrete if call data reveal which rules people cannot understand.

Miyoshi’s publication does not yet establish all three gains. It does something more foundational: it makes them separable. Availability is not accuracy. Accuracy is not resolution. Resolution is not satisfaction. A low transfer count is not the same as a labor saving. A company’s product narrative is not a municipal finding.

The city’s five unsupported answers draw the safety boundary. Its after-hours calls show the service opportunity. Its 61 unscored answers reveal the unfinished information work. Its refusal to claim a 95% workload reduction preserves the value of the experiment.

The next version may answer more questions. The more important achievement would be knowing, with greater precision, which questions it should not answer at all.

Research and sources

Editor’s note: Miyoshi City’s webpage and official 11-page report are the primary basis for findings about the trial. Verbex materials are identified as company statements. Quality review covered a 191-call sample, not all 394 calls, and each metric excludes different unassessable cases. The city webpage reports 195 calls outside its broader opening-hours definition, or 49.5%; the detailed report reports 212 calls outside a narrower “weekday daytime” definition, or 53.8%. Both definitions are disclosed. The supplied exchange-rate timestamp of August 23 at 3:06 p.m. UTC converts to August 24 at 12:06 a.m. JST. Recommendations about future service design are Japan.co.jp analysis based on the disclosed evidence.