In a sixth-grade Japanese question, pupils were asked to use a passage about how white and black clothing absorb heat, quote the source correctly and then explain which color might feel cooler under a fierce sun. Many could reach an opinion. Far fewer could build the bridge from the quoted evidence to their own conclusion. Only 33.6 percent satisfied the task’s conditions.
That bridge appears again and again in Japan’s 2026 National Assessment of Academic Ability and Learning: between a word and its effect, between a diagram and a proportion, between an English opinion and the reason supporting it. The weak point was often not silence or ignorance. It was the unfinished sentence after “because.”
The Ministry of Education, Culture, Sports, Science and Technology and the National Institute for Educational Policy Research released national averages on July 16 and the deeper national analysis on August 3. The test covered sixth-graders in elementary school and third-year students in junior high—children near the end of Japan’s two compulsory-school stages. Roughly 1.8 million pupils were eligible across more than 28,000 schools and related institutions.
Five numbers—and the limits of a league table
Elementary pupils averaged 7.9 correct answers out of 13 in Japanese and 9.1 out of 16 in mathematics. Junior-high students averaged 9.6 out of 15 in Japanese and 9.2 out of 16 in mathematics. English, administered fully by computer, was reported on an item response theory scale: the national four-skill score was 499, with the scale anchored around 500; the three-skill figure excluding the nationally sampled speaking component was 502.
| Assessment | National result | What it represents |
|---|---|---|
| Elementary Japanese | 61.1% | 7.9 correct of 13 questions on average |
| Elementary mathematics | 56.6% | 9.1 correct of 16 questions on average |
| Junior-high Japanese | 64.2% | 9.6 correct of 15 questions on average |
| Junior-high mathematics | 57.4% | 9.2 correct of 16 questions on average |
| Junior-high English | IRT 499 | Listening, reading, writing and speaking combined; speaking national value based on 500 same-day schools |
These are not marks against a fixed annual exam. The content and difficulty vary. The ministry therefore warns against treating a movement in raw percentage as an exact rise or fall in underlying national ability. English’s IRT method is part of the response: it estimates proficiency from the pattern and difficulty of items instead of assuming that every pupil must answer the same paper in the same order.
The autumn release will add prefecture and designated-city analysis. That is when the public temptation to rank places becomes strongest. From the first nationwide test in 2007, however, the ministry has said the exercise measures only a particular portion of academic ability and must not lead to excessive competition or simplistic ordering of schools.
The missing “because”
The official analysis makes the score report vivid by examining errors. In elementary Japanese, the source-based writing task produced a 33.6 percent correct rate. Some pupils expressed a plausible opinion about cool colors but did not quote the collected information or distinguish the quotation from their own idea. The lesson is not simply “write more.” It is to make evidence and inference visibly different.
Junior-high Japanese posed a parallel challenge. Students compared examples using the mimetic word kippari—roughly “firmly” or “decisively”—and explained both the specific manner it conveyed and the broader effect of mimetic language. The correct rate was 39.3 percent. Many could imagine confidence but did not ground the explanation sufficiently in the text they had read.
English exposed the same structure. On a writing task about internet use, 53.5 percent met the conditions, while 43.6 percent gave an opinion without the required reason. On a speaking task asking what a pupil would like to do with an exchange student and why, the correct rate was 28.4 percent. Children often named the activity but omitted or could not organize the reason. A comparable 2019 speaking item had a 45.9 percent correct rate, although the differing tasks mean that figure is context, not a direct trend line.
- Japanese: Quote the source, distinguish it from your idea, then show how the evidence supports the conclusion.
- Mathematics: Identify the quantities or geometric conditions, represent their relationship, then justify the answer.
- English: State an opinion or intended action, add the reason, and organize both so a listener or reader can follow.
Eight cubed, not eight squared
In mathematics, one elementary question asked for the expression giving the volume of a cube with eight-centimeter sides. The answer was 8 × 8 × 8. Yet 10.2 percent wrote 8 × 8, using the area formula for a square, while another 5.5 percent used expressions such as 8 × 6 or 8 × 3. The error suggests that some pupils remembered numbers associated with a cube but had not connected the three-dimensional object to the operation.
A proportion item asked pupils to select diagrams showing a red tape 1.5 times as long as a white one. The correct rate was 34.5 percent. A large wrong-answer group appeared to interpret “1.5 times” as adding one and a half scale divisions, confusing multiplicative comparison with an additive amount. The problem is elementary, but the conceptual fault line runs through later mathematics, science, finance and statistics.
At junior high, only 49.6 percent identified the correct angle relationship needed to demonstrate that two lines were parallel. Many chose angles that looked paired without recognizing the precise corresponding or alternate relationship. The ministry’s prescription is practical: teach geometry through concrete use, and use diagrams to make quantities and ratios visible rather than leaving formulas detached from meaning.
A computer test that listens
The 2026 English assessment marks a technological turn. Listening, reading, writing and speaking were all delivered through computer-based testing. Pupils used screens, audio and microphones over staggered windows from April 20 into May. The three non-speaking skills were administered across all participating schools; the national four-skill estimate used speaking data from 500 schools conducting that portion on the designated schedule.
This is more than substituting a keyboard for a pencil. A speaking assessment can preserve a pupil’s voice, allow common prompts and measure a skill that a paper booklet cannot capture. IRT allows different item sets to be placed on a common scale. It also demands reliable devices, headphones, networks, accommodations, security and support in thousands of schools.
Japan plans to move all subjects in the national assessment to CBT from fiscal 2027. The old image of a nation sitting the same booklet on one April morning is giving way to testing windows and calibrated item banks. That may permit faster, more detailed feedback. It also makes the quality of school digital infrastructure part of the measurement environment itself.
Two digital lives: creating and consuming
The pupil questionnaire draws a sharp distinction between confident use of technology for learning and prolonged consumption outside school. Pupils who said they could organize information, prepare presentations, compose text or gather information with digital tools tended to post higher subject results. About 90 percent also said ICT helped them learn enjoyably and cooperate with classmates.
By contrast, longer weekday use of smartphones for social media and video viewing was associated with lower results in every tested subject. Among elementary pupils, 21.7 percent reported four hours or more a day; among junior-high students, the share was 28.5 percent. In junior-high English, for example, the average IRT score rose from 463 among the four-hours-or-more group to 548 among those reporting less than 30 minutes. That is a striking gradient—but not an experiment.
Heavy screen use may displace sleep, reading or study. It may also reflect stress, household circumstances or disengagement that independently affects learning. Self-reported time is imprecise. The responsible conclusion is not “phones caused the score,” but that the association is large enough for families, schools and researchers to investigate without confusing educational creation with passive consumption.
The quiet architecture of a classroom
Approximately four in five pupils said they had engaged in the “proactive, interactive and deep learning” emphasized by Japan’s curriculum. Those who said they tackled problems on their own initiative tended to score higher. Pupils confident in ICT, pupils who enjoyed reading and pupils who said lessons were easy to understand also tended to perform better.
One finding reaches beyond test performance. Pupils who felt that their teacher recognized their good qualities were more likely to believe that they could express a different opinion in class without being rejected. The correlation was 0.316 in elementary school and 0.331 in junior high. That does not prove a one-directional effect, but it describes something every good classroom needs: intellectual risk feels safer when a child believes the teacher sees value in them.
The survey also found little large gender gap in mathematics performance while boys were more likely than girls to call themselves good at mathematics. Confidence and achievement were not the same thing. Girls led boys substantially in Japanese and in the English IRT result, but the mathematics finding warns against allowing self-perception to harden into expectations before actual ability warrants it.
From the “PISA shock” to a national mirror
The assessment began in 2007 after Japan’s results in PISA 2003 and TIMSS 2003 stirred concern about declining achievement and motivation. The 2005 Basic Policies for Economic and Fiscal Management called for a nationwide survey, and the Central Council for Education endorsed it as part of a mechanism to assure the quality and equality of compulsory education.
The first test separated knowledge questions from application questions in Japanese and mathematics. It was never meant to be a graduation examination. It was a diagnostic mirror: national government would inspect policy, local boards would compare their circumstances with the national picture, schools would revise teaching, and pupils would receive individual feedback.
2007: The modern national assessment begins as a census of sixth- and ninth-graders.
2010 and 2012: The main survey uses sampling plus voluntary participation; 2011 is set aside after the Great East Japan Earthquake.
2013: Full-cohort testing returns; from 2014 it continues as the general format.
2019: Junior-high English joins, including computer-based speaking.
2020: The survey is cancelled because COVID-19 school closures disrupt classrooms.
2025: Junior-high science and pupil questionnaires advance the CBT transition.
2026: All four English skills move online.
2027: All subjects are scheduled for CBT.
The interruptions matter. The Great East Japan Earthquake prevented the planned 2011 survey, and the pandemic eliminated the 2020 administration. They are reminders that an education dataset is never detached from the society producing it. The children taking the 2026 test began elementary school around the pandemic’s onset. Their schooling has spanned closures, masking, accelerated device distribution and the normalization of digital lessons.
Not PISA, not an entrance examination
Japan performs strongly in international studies, but this national assessment answers a different question. PISA samples 15-year-olds and emphasizes applying knowledge across systems; TIMSS samples grades and compares mathematics and science internationally. Japan’s national survey is curriculum-linked, reaches essentially every public school in the target grades, returns information to schools and pairs test answers with detailed pupil and school questionnaires.
Nor should it become a proxy entrance examination. It is administered to pupils, but its stated purpose is system improvement. Once an average becomes a badge of prefectural pride or school reputation, teachers face pressure to teach the test, narrow the curriculum or avoid difficult pupils. The ministry’s warning against rankings is not ceremonial; it protects the diagnostic value of the instrument.
Participation also requires nuance. Public-school implementation was nearly universal, while private-school participation was far lower—about half of eligible elementary private schools and roughly a quarter of junior-high private schools. National figures combine participating national, public and private institutions, but the assessment is most nearly comprehensive for the public system.
What teachers can take back in September
A score release is useful only if it returns to the lesson. For language teachers, the 2026 work is to make evidence visible: highlight the quoted phrase, name the inference and ask whether the second truly follows from the first. For mathematics teachers, it is to move repeatedly between object, diagram, words and equation. For English teachers, it is to make “opinion plus reason” habitual in short exchanges before demanding it under test conditions.
The official response includes analysis tools for education boards, detailed reports, lesson ideas and research access to microdata. It also reaches beyond pedagogy: smaller junior-high classes, subject-specialist teaching in elementary grades, support staff, teacher-workload reform and school environments that improve well-being. A child’s explanation cannot be improved by a worksheet alone if a teacher has no time to read it.
Local analysis due in autumn will show variation among prefectures and designated cities. The best question will not be “Who won?” It will be “Which wrong answers recur here, which groups are being missed, and what can tomorrow’s classroom do differently?”
The national portrait inside one sentence
Japan’s assessment began as a response to anxiety about national standing. Nineteen years later, its most useful contribution is more intimate. It reveals the moment a pupil has an idea but cannot yet make it accountable to evidence; knows a formula but cannot see the shape inside it; can speak an intention but loses the reason on the way to the microphone.
Those are not failures to be pinned to a child, teacher or prefecture. They are precise places where teaching can begin. The 2026 report also shows resources already present in classrooms: pupils willing to collaborate, confidence with creative digital tools, affection for reading, and the powerful sense that a teacher recognizes what is good in them.
The five national averages will lead the headlines. The deeper portrait is of a system trying to move from recall to explanation, from paper uniformity to calibrated digital assessment, and from ranking schools to understanding learning. Japan’s classrooms do not need only more correct answers. They need more complete reasons—and enough trust for children to say them aloud.
Reporting notes and principal sources
National averages, item analyses and questionnaire relationships are from MEXT and the National Institute for Educational Policy Research. Correlations are reported as associations and do not establish causation.
- MEXT: 2026 assessment reports and three-stage release schedule
- NIER: Complete 2026 reports and result materials
- MEXT/NIER: Summary analysis points, August 2026
- NIER: Elementary participation and national results
- NIER: Junior-high participation and national results
- MEXT: Purpose and history of the national assessment
- MEXT: 2007 guidance on use and the warning against rankings
- MEXT: Cancellation of the 2020 assessment
- MEXT: National-assessment CBT transition
