Large-scale genomics has traditionally rewarded one architecture: put a large storage system beside a large compute cluster and bring the data to the machines. Mitsui Knowledge Industry and NTT DOCOMO BUSINESS are testing whether that assumption can be loosened.
On Oct. 5, the companies announced results from a joint demonstration using NTT DOCOMO BUSINESS’s “GPU over APN Testbed.” About 5 terabytes of genomic variant data—roughly 17 billion records from 3,202 individuals—were stored in the Tokyo metropolitan area while GPU workloads were distributed across five geographically separated sites from Sapporo to Fukuoka. Across three representative genomic queries, processing time fell to 25–29% of the single-site baseline.
Five sites, one 5-terabyte genomic workload
The data consisted of Variant Call Format records from 3,202 publicly available genomes. The underlying collection traces to the 1000 Genomes Project sample set and the later high-coverage sequencing program coordinated through the New York Genome Center and IGSR.
The companies ran three use cases: finding samples with variants in a specified chromosome region; counting samples with a genotype associated with known disease risk; and identifying samples carrying any of 220 known cancer-related driver variants.
From 134 seconds to 33
For the first query, one-site processing took 134.4 seconds; five-site processing took 33.5 seconds. The second fell from 99.9 to 28.6 seconds, and the third from 102.3 to 28.9 seconds. Five locations did not produce perfect fivefold scaling, but the real-world-style workload ran roughly three-and-a-half to four times faster.
| Use case | One site | Five sites | Ratio |
|---|---|---|---|
| Variant search in chromosome region | 134.4 sec | 33.5 sec | 25% |
| Disease-risk genotype count | 99.9 sec | 28.6 sec | 29% |
| 220 driver-variant search | 102.3 sec | 28.9 sec | 28% |
The gain was not simply a matter of adding GPUs. Mitsui Knowledge Industry contributed workload-optimization techniques, including storage-access optimization, and high-speed genomic search technology. In data-intensive science, adding remote compute helps only if storage access, partitioning and network transfer do not become the new bottleneck.
Why the network becomes part of the computer
NTT DOCOMO BUSINESS launched GPU over APN Testbed in July 2026. It places GPUs at eight sites—Sapporo, Kanazawa, Osaka, Fukuoka and four locations in the Tokyo metropolitan area—and connects them through NTT’s IOWN All-Photonics Network. The company describes 100-Gbps-class low-latency, high-capacity links, with virtual-machine and Kubernetes support and multi-tenant operation.
The idea responds to a physical constraint in modern computing. Demand for GPUs has surged with generative AI, but data centers are constrained not only by chip supply. Power delivery, cooling, rack capacity and local grid limits can all prevent operators from concentrating enough accelerators in one building.
Distributed computing makes the network part of the computing system. If distant GPU pools can behave sufficiently like one resource, capacity can be placed where electricity, space or hardware is available rather than where one data center happens to have room.
Genomics is a demanding test case
Genome analysis is not just arithmetic. Many workloads are dominated by storage I/O, indexing, filtering and movement of very large files. As cohorts grow and sequencing becomes deeper, the cost of moving and searching data can rival the cost of computation itself.
That makes the NTT–Mitsui test more informative than a synthetic GPU benchmark. The companies ran actual genomic-search use cases and kept the dataset centralized while dispatching compute across the country. Architecturally, that raises a fundamental question for scientific computing: should data move to computation, or should computation move toward data?
The public-data lineage runs back to the 1000 Genomes Project
The 3,202-sample dataset exists because of a long shift toward shared genomic reference resources. The international 1000 Genomes Project, launched in 2008, created a widely accessible map of human genetic variation. Its 2015 Phase 3 analysis centered on 2,504 individuals. The New York Genome Center later produced approximately 30x high-coverage sequence data for 3,202 samples, including additional related individuals, with support from the U.S. National Human Genome Research Institute.
Public datasets are ideal for infrastructure testing because access and reuse are already defined. Clinical genomes introduce a different class of requirements: consent, privacy, access control, auditability, institutional governance and, potentially, restrictions on where data may be processed.
Japan’s genomic-medicine framework is also maturing
Japan enacted its Genome Medicine Promotion Act in June 2023, formally titled the Act on Comprehensive and Systematic Promotion of High-Quality and Appropriate Genomic Medicine That Citizens Can Receive with Confidence. A national basic plan followed in November 2025.
As genomic medicine expands, the bottleneck moves beyond sequencing machines. Storage, search, compute, networks and data governance all become part of the clinical and research infrastructure.
Could hospitals one day tap unused GPUs around Japan?
The two companies say the knowledge from the demonstration will be applied to data-intensive research fields and could support hospitals, universities and research institutions. NTT DOCOMO BUSINESS also plans to apply distributed-GPU technology to GPU-as-a-Service offerings and its broader AI-Centric ICT platform.
That future is not yet a production medical service. Real patient data would require decisions on where information may reside, whether raw data can cross institutional boundaries, how encryption and authentication work, how audits are maintained and what happens when a site or network path fails.
- Scaling efficiency beyond five GPU sites
- Performance when genomic datasets grow far beyond 5 TB
- How heterogeneous GPU models and site performance are scheduled
- Security and governance requirements for real patient data
- GPUaaS pricing compared with dedicated on-premises clusters
- Resilience during site, power or network failures
An alternative to building one ever-larger machine
Generative AI has turned GPUs from specialized hardware into strategic infrastructure. The question of where to place them now touches electricity, land, cooling, networking and regional resilience.
The NTT–Mitsui demonstration does not prove that every scientific workload should be distributed. It does show that, for these genomic-search tasks, five geographically separated GPU sites can materially outperform a single-site execution while working against a centrally stored dataset.
The harder question comes next: can this become cheaper, simpler, secure and reproducible enough for everyday researchers? If so, access to genomic computing may gradually depend less on which institution owns the largest cluster—and more on whether researchers can reach a national pool of compute when they need it.
Sources & Reference Material
- Mitsui Knowledge Industry and NTT DOCOMO BUSINESS, nationwide distributed-GPU genomic-analysis demonstration, Oct. 5, 2026.
- NTT DOCOMO BUSINESS, launch of GPU over APN Testbed, July 6, 2026.
- Japan Ministry of Health, Labour and Welfare, genomic-medicine framework and Genome Medicine Promotion Act.
- International Genome Sample Resource, 3,202 high-coverage samples from NYGC, Aug. 14, 2020.
- NTT and NTT DOCOMO, remote-GPU In-Network Computing demonstration over IOWN APN, March 2, 2026.
Reporting cutoff: Oct. 6, 2026, 1:21 PM JST. Public materials reviewed do not specify the exact GPU models and counts, per-site hardware configuration, performance on real patient data, a production security architecture, GPUaaS pricing, or scaling beyond five sites.
