When your carrier boasts about having „the best network,“ that claim almost always traces back to one of a handful of independent testing houses. Three names dominate the conversation — Umlaut, Ookla, and P3 — and although they are often mentioned in the same breath, they measure networks in fundamentally different ways. Understanding those differences is the difference between reading a marketing badge and actually knowing what it means.
First, the plot twist: they are more connected than you think
Before comparing methodologies, it helps to untangle the corporate history, because it explains a lot of the confusion. „P3“ and „umlaut“ are not really two competitors — they are the same lineage. P3 communications built its reputation over two decades as the benchmarking arm of the P3 group, and its „P3 Score“ became a de facto industry standard applied across more than 80 countries. In October 2019, P3 group rebranded the entire company as umlaut, and the familiar „P3 connect Mobile Benchmark“ became the „umlaut connect“ benchmark.
It went further. In June 2021, Accenture acquired umlaut, folding it into its engineering services division — which is why recent reports carry the „connect“ branding under the Accenture umbrella. And in a twist almost nobody saw coming, in March 2026 Accenture agreed to acquire Ookla too — the company behind Speedtest, Downdetector, RootMetrics and Ekahau — in a $1.2 billion deal. In other words, the two dominant and philosophically opposed schools of network measurement are now heading under the same roof. For two decades Speedtest functioned as a quasi-neutral referee; that neutrality is now a live question. All the more reason to understand how each method actually works.
Umlaut: the controlled, holistic benchmark
Umlaut’s philosophy is to measure the maximum capability of a network under carefully controlled conditions. Rather than relying purely on whatever data happens to trickle in from users, umlaut sends out teams to do the testing itself.
The core of the method is drive testing and walk testing. Fleets of measurement cars — often four or more in a single country campaign — cover cities, towns and highways, while walk-test teams cover pedestrian zones, indoor locations and even train routes between cities. Every vehicle and walker carries the latest smartphones running scripted, repeatable tests on 4G and 5G. This produces two headline categories: voice and data. Voice testing measures call setup success, call setup time, and speech quality using the POLQA wideband algorithm, while checking whether data remains available during a call. Data testing measures download and upload throughput, web page loading (both national and international sites), file transfers using fixed-size files, and video streaming quality.
On top of the controlled tests, umlaut layers a crowdsourcing component — the „User Experience Test“ — drawing on billions of background samples collected from real devices over many weeks. This measures broadband coverage, real-world download speeds, latency and time spent connected to 4G and 5G. The results are combined into the umlaut Score on a 0–1000 scale, translated into school-style grades (outstanding, very good, good, and so on). In its European benchmarks, the controlled voice and data tests carry most of the weight, with crowdsourcing contributing roughly a quarter of the final score. Umlaut publishes its full method in its Telecommunications Benchmarking hub and the detailed 2025 Mobile Benchmarking Framework white paper.
Ookla: the crowdsourced speed referee
Ookla takes the opposite approach. Its data comes almost entirely from consumer-initiated Speedtests — the millions of people who open the app or website to check their connection. That produces enormous scale: more than 250 million tests per month, capturing over 1,000 attributes per test across real devices, locations and network conditions.
Each test measures three fundamentals: download speed, upload speed, and latency (ping). For its rankings, Ookla distills these into proprietary metrics. The Speed Score combines download and upload speeds — historically weighted heavily toward download, at roughly a 9:1 ratio — with newer awards also factoring in loaded latency. The Consistency Score measures the share of samples that clear a minimum speed threshold, rewarding networks that are reliably fast rather than occasionally spectacular. For fixed broadband, the Connectivity Score blends raw speed (50%) with video streaming (25%) and web browsing experience (25%). To even qualify for a Speedtest Award, a provider must account for at least 3% of samples in a country and cannot be an MVNO.
The strength here is authenticity and scale: these are the speeds real people experienced on their own phones. The weakness is bias — tests cluster where users already have good signal and choose to run them, so pure crowdsourcing can under-represent dead zones and rural gaps. Ookla documents exactly how the awards are calculated in its Speedtest Awards methodology.
P3: the methodology that started it all
If umlaut’s method feels rigorous, that is because it is the P3 method — inherited wholesale. P3 pioneered the combination of drive testing, walk testing, railway testing and crowdsourcing into a single holistic score. When people still refer to „P3 testing,“ they are almost always describing the same drive-and-walk-plus-crowd approach now branded umlaut (and, most recently, simply „connect“ under Accenture). Treat P3 as the historical name for the benchmark, not a separate living methodology. The lineage is confirmed in the official P3-becomes-umlaut announcement.
Side-by-side comparison
| Attribute | Umlaut (connect) | Ookla (Speedtest) | P3 (legacy) |
|---|---|---|---|
| What it is | Independent holistic benchmark, now under Accenture | Crowdsourced speed measurement platform | Original benchmark; rebranded to umlaut in 2019 |
| Primary focus | Maximum network capability plus real-world experience | Real-world consumer-experienced speed and reliability | Same holistic focus as umlaut (its predecessor) |
| Data collection | Drive tests, walk tests, railway tests + crowdsourcing | Consumer-initiated app/web tests (250M+/month) | Drive, walk, railway tests + crowdsourcing |
| Conditions | Controlled, scripted, repeatable | Uncontrolled, organic, user-triggered | Controlled, scripted, repeatable |
| Evaluation criteria | Voice (call success, setup time, POLQA speech quality); Data (download/upload, web, file transfer, video); Crowd (coverage, latency, 4G/5G time) | Download speed, upload speed, latency, jitter; consistency vs. thresholds; video & browsing for fixed | Voice, data and crowd KPIs — identical framework to umlaut |
| Scoring model | umlaut Score, 0–1000, mapped to school grades | Speed Score, Consistency Score, Connectivity Score | P3 Score (historical, 0–1000 basis) |
| Coverage of dead zones | Strong — teams actively drive/walk weak areas | Weaker — samples cluster where users have signal | Strong — same active testing model |
| Video experience | Measured inside controlled data tests (streaming quality, start time, stalling) | Dedicated Video Score; video weighted 25% in the fixed Connectivity Score | Same as umlaut — video within data testing |
| Gaming / real-time response | Captured via general latency KPIs; no dedicated gaming score | Strongest — dedicated Game Score using median game latency, jitter, download & upload, measured to real game servers | Latency KPIs only (legacy) |
| AI / agentic response time | No dedicated metric; latency and consistency serve as proxies | No dedicated metric yet, but latency-first focus + Accenture’s „agentic, low-latency“ direction lean this way | No dedicated metric |
| Best for | Operators, regulators, „best network“ certification | Consumers, marketers, fast broad snapshots | Historical reports and legacy references |
What about video, gaming, and AI response times?
This is where the two philosophies diverge most sharply. Ookla has leaned hard into experience-specific scores: a dedicated Video Score for streaming quality and a Game Score built specifically around responsiveness — median game latency and jitter measured to real game servers, blended with download and upload. Because gaming and video are about how quickly and consistently packets arrive, and Ookla’s whole model is latency-and-consistency-first, it is the natural home for these metrics.
Umlaut (and therefore legacy P3) does test video streaming, but as one component inside its controlled data suite rather than as a headline consumer score, and it has no public gaming-specific rating — its edge is breadth and rigor of coverage, not experience-tier granularity.
And AI response time? It is worth being precise here: none of these network benchmarks yet publishes a dedicated „AI responsiveness“ score. Metrics like time-to-first-token belong to LLM benchmarking, not network testing. What the network testers do measure — latency, jitter and consistency — is exactly what determines how snappy an AI assistant or agent feels over a mobile connection. Given that Accenture framed its Ookla acquisition explicitly around „agentic access“ and „low-latency, zero-friction connectivity,“ expect an AI-responsiveness metric to emerge from the Ookla/umlaut camp before long. For now, treat latency and consistency as the honest proxies.
So which one should you trust?
It depends on the question you are asking. If you want to know how good a network can be — including in the places people rarely test — umlaut’s controlled drive-and-walk approach is hard to beat, which is why operators and regulators lean on it. If you want a fast, honest snapshot of what typical users are actually getting today, Ookla’s crowdsourced Speed and Consistency Scores are the most representative of everyday experience. And P3? It is the historical root of the entire controlled-testing school — worth knowing when you stumble across an older report.
The most interesting development is that this distinction may soon blur. With both umlaut and Ookla converging under Accenture, the long-standing tension between controlled benchmarking and crowdsourced measurement could turn into a single combined offering — powerful for operators, but a reason for the rest of us to keep asking exactly how any „best network“ badge was earned.
Sources & further reading
- Umlaut / connect: Accenture Telecommunications Benchmarking hub, and the 2025 umlaut Mobile Benchmarking Framework (white paper, PDF).
- Ookla: Speedtest Awards Methodology (2025) — covering Speed Score, Consistency Score, Video and Game Scores.
- P3: P3 group AG becomes umlaut (2019 rebrand announcement).

