Identity and terminology note: The documented company is AI Native of Shibuya, incorporated in October 2025 with stated capital of ¥1 million. It has not published an employee count, so “tiny” should not be read as a verified headcount. Its June announcement uses the shorthand “engineers and PMs,” but the detailed audience names PdM—product managers—alongside engineers, tech leads and CTO/VPoE roles. The supplied headline says project managers. This article preserves that headline while keeping the two professions distinct.

The first question begins with a feature request. When a new function lands on your desk, how do you divide the code and separate responsibilities? The choices form a small staircase: prioritize code that runs and tidy the structure later; divide modules deliberately and explain the design; or establish an architecture, guidelines and automated checks that make quality repeatable.

No compiler watches the answer. No colleague opens the pull request. The respondent recognizes—or chooses—the description that feels most like their work. Then the form advances. Twenty questions later, a radar chart gives shape to an identity that is usually scattered across repositories, incident calls, meeting notes and the memories of teammates.

This is AI Native’s “AI Development Skills Assessment,” publicly announced on June 22. The Shibuya company says it can visualize an engineer’s strengths and room to grow across nine axes. The news release promises roughly seven minutes; the current site says 21 questions and about ten. The change may be ordinary product copy drift, but it is also a reminder that the instrument itself is moving.

The product arrives at a useful moment. Generative AI made it dramatically easier to produce a convincing prototype. That moved the bottleneck. The valuable worker is no longer merely the person who can make a model answer. It is the team that can decide whether AI belongs in a problem, specify its behavior, evaluate failures, deploy it, watch it change and stop it from leaking or harming anything.

21 questionsThe length of the current public self-diagnostic
9 axesThree foundations and six AI-specific domains
2 viewsCapability and experience are shown separately
¥1 millionAI Native’s stated capital at incorporation

A startup young enough to count in months

AI Native says it was established in October 2025. Its registered address is in Maruyamacho, a steep-edged Shibuya neighborhood where small offices sit between hotels, music venues and the roads flowing toward Dogenzaka. By the assessment’s release the following June, the corporation was about eight months old.

The company lists AI media, generative-AI products, consulting and development as its businesses. The assessment is one of three free diagnostic lines: a broad AI-use check for employees, an organizational maturity framework for executives, and this development-focused form. Free diagnostics can educate users, attract consulting leads and reveal what language prospective clients use for their problems. Those purposes can coexist.

Its youth makes the breadth of the scale striking. The form does not ask only whether someone can prompt a model. It asks the organization to think from problem definition through operations. In the company’s framework, three foundational axes—engineering basics, verbalization and communication, and efficiency or return-on-investment thinking—support six specialized ones.

AxisWhat AI Native says it coversEvidence a mature assessment could seek
Engineering foundationsDesign across layers and testing strategyCode, architecture decisions, reviews and test behavior
CommunicationStructured thought, instructions to AI and agreement-buildingRequirements, review discussions and stakeholder outcomes
Efficiency and ROIInvestment logic, waste removal and improvement cyclesBaseline cost, adoption, quality and realized benefit
Problem and requirements definitionDeciding where AI fits and translating a need into requirementsRejected as well as accepted use cases; measurable acceptance criteria
AI-domain knowledgeContext engineering and multi-agent conceptsDesign choices, failure analysis and technically accurate explanation
Quality assuranceEvaluation design and evaluation pipelinesTest sets, metrics, adversarial cases, regressions and release gates
OperationsLLMOps, deployment and model updatesLogs, alerts, rollback, cost controls and incident response
Process designRedesigning development around AIChanged handoffs, review ownership and measured cycle time
Security and governanceGuardrails and incident responseThreat models, access controls, audits and exercised playbooks

The third column is not part of AI Native’s published product. It shows the evidence that would turn a vocabulary into a defensible organizational assessment. That distinction is central. A list can name the right things without yet proving that its score measures them.

What the twenty-one questions actually measure

The site calls the result objective in places, but the live form is structured self-report. The first item offers three polished descriptions of behavior, and a hint tells respondents to judge whether they could explain the architecture of a recent pull request. That is more concrete than asking “Are you good at design?” It anchors the choice in remembered work. Still, the work itself is not observed.

AI Native separates “Capability,” described as knowledge or design ability, from “Experience.” This is an intelligent choice. Someone can explain an evaluation pipeline without having operated one during a midnight incident. Another engineer may have deployed several systems through a vendor platform without being able to generalize the design. Knowledge and experience should not collapse into one number.

Yet 21 items distributed across nine axes leave little redundancy. Reliability normally benefits from several independent items that probe the same construct from different angles. A single attractive answer can reflect competence, aspiration, generous self-interpretation or simple familiarity with current vocabulary. The public material does not disclose its scoring weights, item-development process, reliability, test-retest results, validation sample or comparisons with observed job performance.

That absence does not make the tool worthless. It places it in the correct category: a brief developmental conversation starter. Research synthesizing 22 meta-analyses found an average correlation of .29 between self-evaluated ability and objective performance—meaningful but modest, and better when questions are specific and familiar. AI Native’s concrete prompts may help. They do not erase the limitation.

A radar chart can show where to ask the next question. It cannot, by itself, answer whether someone can safely ship an AI system.

A better second stage would triangulate the result. Ask for an architecture decision, a small evaluation design, an incident narrative and a work sample. Let a manager or peer rate demonstrated behavior against the same anchors. Compare the sources, investigate gaps and repeat after real work. The score then becomes the beginning of evidence rather than a substitute for it.

The denominator hiding inside the percentages

AI Native’s launch release includes an alluring finding. Among early users from March through June, the share rated “High” reached 88.6 percent in efficiency and ROI thinking, 85.7 percent in AI-domain knowledge, and 82.9 percent in communication and problem definition. “Low” appeared most often in LLMOps at 17.1 percent and quality assurance and engineering foundations at 14.3 percent.

The pattern tells an appealing industry story: people know how to propose AI, but fewer can keep it reliable in production. That story is plausible. NIST’s AI Risk Management Framework and its generative-AI profile devote substantial attention to evaluation, monitoring, governance and risk across the lifecycle precisely because a demonstration is not a dependable system.

But the press release does not state the sample size or recruitment method. The decimals contain a clue. The full set of percentages is exactly consistent with a denominator of 35: 31 of 35 is 88.6 percent when rounded; 30 is 85.7; 29 is 82.9; six is 17.1; and five is 14.3. Multiples of 35 could produce the same rounded pattern, so 35 is an inference, not a reported fact. For a new tool’s “advance users,” it is the parsimonious reading.

Published percentagePossible count if n = 35What can safely be concluded
88.6% High: efficiency and ROI31High self-ratings were common in this early group
85.7% High: AI-domain knowledge30The recruited group considered itself AI-literate
82.9% High: communication and requirements29Upstream confidence was common
17.1% Low: operations and LLMOps6A minority selected low operational anchors
14.3% Low: QA and engineering basics5 eachSome respondents identified production-quality gaps

Nothing here estimates Japan’s engineering workforce. Early users of a specialist startup’s diagnostic are self-selected; they may be unusually engaged with AI, close to the company, or motivated to present themselves well. “High” and “Low” are categories generated by the vendor’s rubric, not standardized percentiles. The figures are useful as product discovery, not labor statistics.

Five questions the next release should answer
  • How many people completed the assessment, and how were they recruited?
  • How were the items and cut scores written, reviewed and revised?
  • Do scores remain stable when the same person retakes the form?
  • How do results relate to work samples, peer ratings and project outcomes?
  • Do items perform differently by role, experience, gender, language or company type?

Before dashboards, the long attempt to put skill into boxes

The desire to measure work is older than software. Industrial management split jobs into observable motions and times. Professional certification later tried to identify the knowledge a practitioner should possess. In each era the promise was portability: if skill could be described independently of one manager’s intuition, workers and organizations could compare, teach and plan.

In 1980, Hubert and Stuart Dreyfus proposed a staged account of skill acquisition. Novices rely on context-free rules; experience gradually changes what the practitioner notices, how a situation is organized and how decisions are made. Later versions became widely known through five stages from novice to expert. The durable insight is not the labels. It is that expertise is not merely a larger inventory of facts. It is perception and judgment formed in concrete cases.

Project management followed its own route toward professional definition. The Project Management Institute was born in 1969 after practitioners sought a place to share planning and scheduling problems. Its first PMP examination was held in 1984: 56 people sat it and 43 passed. A profession that had been scattered across engineering, construction, defense and business gained a credential and common body of knowledge.

Digital work then required a map broad enough to outlive particular products. SFIA, formally launched in 2000 from collaborative initiatives dating to the 1980s, arranged professional skills against seven levels of responsibility. Its current language runs from “Follow” to “Set strategy, inspire, mobilise.” Crucially, SFIA says competence is demonstrated in real work: a qualification that tests knowledge alone does not establish experience or responsibility.

1969 Project Management Institute is formed, helping define project management as a profession.

1980 The Dreyfus brothers publish a staged model of skill acquisition rooted in experience.

1984 PMI administers the first PMP examination.

2000 SFIA is formally launched as a shared language for digital skills and responsibility.

2002 Japan’s Ministry of Economy, Trade and Industry establishes ITSS.

2014 IPA releases the i Competency Dictionary, linking tasks, skills and knowledge.

2022 METI and IPA publish the Digital Skill Standards for economy-wide digital transformation.

2024 DSS 1.2 adds generative-AI material and explicitly discusses product managers.

2026 DSS 2.0 refreshes data, AI and business-transformation roles; AI Native releases its rapid diagnostic.

These frameworks solve different problems. Dreyfus is a theory of learning. PMP is a professional credential. SFIA is a reference language for responsibility and skill. A startup diagnostic is a fast interface for reflection and commercial service discovery. Trouble begins when one kind is treated as another.

Japan’s road from ITSS to DSS 2.0

Japan already has an unusually explicit public history of digital skill scales. METI created the IT Skill Standard in 2002 as the industry worried about global competition, offshore development and the cultivation of high-level professionals. The framework did not simply label everyone “systems engineer” or “programmer.” Its later form defined 11 occupational families, 35 specialties and seven levels based on ability and achievement, drawing in part on SFIA.

ITSS included project management alongside consultants and IT specialists. It tried to connect career paths, experience and training. IPA’s own guidance warned that a skill standard was only a ruler: without a strategy that assembled people’s abilities into useful services, individual scores would not create competitiveness. That caution reads as if written for the radar-chart era.

The 2014 i Competency Dictionary changed the angle. It organized both the tasks an IT-enabled business performs and the skills and knowledge that support those tasks. A company could begin with strategy and work, then design development around the gap. It was less a universal player rating than a configurable dictionary.

Digital transformation expanded the audience beyond IT vendors. METI and IPA issued the Digital Skill Standards in December 2022: a literacy standard for all businesspeople and a professional standard for people driving transformation. It defined software engineers alongside business architects, designers, data scientists and cybersecurity professionals.

The standard began changing almost immediately. Generative AI entered the literacy material in 2023. Version 1.2 in July 2024 added AI-development examples and a clarification on product managers. In April 2026, only two months before AI Native’s release, DSS 2.0 refreshed data-management, AI and business-transformation roles.

This fast revision cycle explains the appeal of a startup’s lighter instrument. A national standard is broad, negotiated and slow enough to remain stable. A 21-question web form can add context engineering, multi-agent systems, evaluation pipelines and LLMOps while those terms are still being argued over. The trade is speed for institutional depth.

Product manager, project manager—and two dangerous letters

“PM” is not a stable job title. A project manager organizes a temporary endeavor: scope, dependencies, schedule, resources, risk, contracts and delivery. A product manager is responsible for the continuing direction of a product: whose problem matters, what outcome is sought, what to build, what to reject and how learning changes the roadmap. In Japanese technology companies, PdM is commonly used to disambiguate the second role.

AI Native’s press-release headline says engineers and PMs. Its detailed tool summary names engineers, tech leads, CTO/VPoE and PdM. The public product page focuses on engineers, technical leads and technology executives. It does not publish a separate project-management scale. Therefore the safest reading is product manager, not every professional project manager.

The overlap is real. Both roles define requirements, negotiate tradeoffs, communicate uncertainty and help teams deliver. AI work blurs the boundary further because model behavior turns product policy into an engineering test. But an assessment that includes module boundaries, evaluation pipelines, LLMOps and guardrails is measuring participation in AI-product development, not the full practice of project management.

This matters in hiring. A certified construction-project manager could be exceptional at procurement, sequencing and stakeholder control yet rate low on context engineering. That would not show low project-management ability. It would show that the instrument measures a narrower domain. Validity always has a sentence after it: valid for what interpretation, for which people, for which decision?

The two letters “PM” can make a niche diagnostic look universal. The detailed audience list restores the boundary: this is a map for AI-product teams.

Why a radar chart feels like a game

Game interfaces make growth visible. A role-playing character gains experience, unlocks abilities and reveals an uneven build: strength high, defense weak, magic waiting. The pleasure lies in turning an uncertain future into a next action. Skill dashboards borrow that grammar. A low spoke is not merely failure; it is a quest.

That can be humane. Conventional performance reviews often hide the rules until promotion time. A shared scale lets a junior engineer ask what “better” looks like and lets a manager distinguish technical breadth from organizational influence. Separating capability from experience prevents a person with strong knowledge but few opportunities from disappearing inside one seniority score.

Visualization also compresses. Two identical radar shapes can conceal different histories. One respondent may have built a safe retrieval system for health information; another may have selected the same self-description after a tutorial. Context engineering for a toy chatbot is not context engineering for a multilingual regulated service. Experience is not a liquid that fills any container.

Scores alter behavior. Once tied to promotion, assignment or pay, people learn the rubric. They collect visible artifacts, adopt fashionable vocabulary and avoid valuable work that the chart ignores. The familiar paraphrase of Goodhart’s law applies: when a measure becomes a target, it stops behaving like the same measure.

The remedy is not to hide the scale. It is to keep judgment plural. Combine the chart with outcomes, work evidence, peer perspectives and the conditions under which the work occurred. Reward learning from incidents, not just an incident-free record. Version the framework and preserve old results so that a person’s score does not appear to fall simply because the ruler changed.

When measurement becomes management

AI Native suggests that organizations could use the diagnostic for hiring, training and assignment. Those are progressively higher-stakes uses. A private reflection can tolerate roughness. A training conversation needs enough consistency to point people toward sensible practice. A hiring or promotion screen must withstand questions about relevance, reliability, access, bias, privacy and appeal.

A 10-minute self-report should not become an automatic gate. Candidates differ in modesty, language, familiarity with the answer style and opportunity to encounter advanced systems. People from organizations with strong governance may accurately report narrow authority; people from informal teams may claim broad ownership because nobody else existed. The score can invert confidence and competence.

For team development, however, the nine axes are a strong meeting agenda. If everyone selects high AI-domain knowledge but low operations, the response is not necessarily to hire an “LLMOps hero.” The team might create a release checklist, define model and prompt versioning, build an evaluation set, assign rollback ownership and rehearse an incident. A scale earns value when it changes the system around people, not merely ranks the people inside it.

Governance must also cover the assessment data. A skill profile is employment information. Organizations need to say who can see it, how long it remains, whether individuals can correct context, and whether a vendor may aggregate it. Email is optional on the public form, according to the site, and browser results are available without entering it. Enterprise use requires a fuller policy.

A responsible organizational workflow
  • Define the decision: reflection, learning, assignment or selection are not interchangeable.
  • Set the evidence: self-report, manager observation, peer input, work samples and outcomes.
  • Calibrate: discuss examples until raters apply anchors consistently.
  • Protect: minimize data, limit access, establish retention and allow correction.
  • Test fairness: inspect score and outcome differences across relevant groups and roles.
  • Review impact: ask whether the process improved learning, quality and opportunity.

For the AI work itself, the scale points in a sound direction. NIST organizes risk practice around Govern, Map, Measure and Manage, and its generative-AI profile emphasizes lifecycle controls. AI Native’s quality, operations and governance spokes translate much of that institutional language into the daily identity of a development team. The missing step is validating the translation.

The test this test still has to pass

A credible assessment needs a claim, evidence and a boundary. AI Native’s claim is that 21 responses can visualize strength and room for growth across nine elements of AI development. Evidence could accumulate in stages: expert review of item content; interviews showing that respondents interpret items as intended; internal-consistency and retest analysis; comparison with work samples and observed behavior; and studies showing whether recommended actions improve results.

The company should publish version numbers and a technical note. It should state the sample behind every percentage, recruitment method, missing-data rules and uncertainty. It should explain whether “High” is criterion-based—meeting a defined behavior—or norm-based—scoring above other people. Those are not cosmetic details. They determine what the picture means.

It should also resist the temptation to claim “market value” too soon. Market value depends on role, sector, geography, language, company stage and a worker’s opportunity to produce results. A radar chart can help articulate capability. It cannot quote a salary without labor-market data and a defensible model.

The most promising feature is the framework’s center of gravity. It places quality assurance, LLMOps and governance beside prompting and domain knowledge. That tells engineers and product managers that production is the profession. A prototype that works once is a demonstration. A system that can be evaluated, monitored, secured, repaired and explained is a service.

There is a scene hidden behind every spoke. Someone writes the first adversarial test. Someone notices a model-cost spike. Someone tells a product leader that the data cannot support the promise. Someone rolls back an update before customers wake. Someone documents the incident so the next person is less alone. Those acts are skill, but also responsibility and organizational permission.

AI Native has made a compact map of that country. The map is not yet a surveyor’s instrument, and its launch data should not be mistaken for a census. Still, it names the mountains many teams discover only after a prototype reaches real users. If the young company now measures its own measure with the rigor it asks of AI systems—tests, versioning, monitoring, transparent failure—it could build something more valuable than a score: a shared language for the work that begins after the demo.

Research notes and principal sources

This article reviewed public information through August 11, 2026 at 8:06 AM JST. AI Native’s early-user figures and product claims are company-published, not independently audited. The possible n = 35 is an arithmetic inference from the published rounded percentages, not a sample size disclosed by the company.