Umlaut, Ookla, and P3: How the Big Three Mobile Network Tests Actually Work

Umlaut, Ookla and P3 all rank mobile networks - but measure them very differently. Compare their focus and evaluation criteria in one clear guide.

When your carrier boasts about having „the best network,“ that claim almost always traces back to one of a handful of independent testing houses. Three names dominate the conversation — Umlaut, Ookla, and P3 — and although they are often mentioned in the same breath, they measure networks in fundamentally different ways. Understanding those differences is the difference between reading a marketing badge and actually knowing what it means.

First, the plot twist: they are more connected than you think

Before comparing methodologies, it helps to untangle the corporate history, because it explains a lot of the confusion. „P3“ and „umlaut“ are not really two competitors — they are the same lineage. P3 communications built its reputation over two decades as the benchmarking arm of the P3 group, and its „P3 Score“ became a de facto industry standard applied across more than 80 countries. In October 2019, P3 group rebranded the entire company as umlaut, and the familiar „P3 connect Mobile Benchmark“ became the „umlaut connect“ benchmark.

It went further. In June 2021, Accenture acquired umlaut, folding it into its engineering services division — which is why recent reports carry the „connect“ branding under the Accenture umbrella. And in a twist almost nobody saw coming, in March 2026 Accenture agreed to acquire Ookla too — the company behind Speedtest, Downdetector, RootMetrics and Ekahau — in a $1.2 billion deal. In other words, the two dominant and philosophically opposed schools of network measurement are now heading under the same roof. For two decades Speedtest functioned as a quasi-neutral referee; that neutrality is now a live question. All the more reason to understand how each method actually works.

Umlaut: the controlled, holistic benchmark

Umlaut’s philosophy is to measure the maximum capability of a network under carefully controlled conditions. Rather than relying purely on whatever data happens to trickle in from users, umlaut sends out teams to do the testing itself.

The core of the method is drive testing and walk testing. Fleets of measurement cars — often four or more in a single country campaign — cover cities, towns and highways, while walk-test teams cover pedestrian zones, indoor locations and even train routes between cities. Every vehicle and walker carries the latest smartphones running scripted, repeatable tests on 4G and 5G. This produces two headline categories: voice and data. Voice testing measures call setup success, call setup time, and speech quality using the POLQA wideband algorithm, while checking whether data remains available during a call. Data testing measures download and upload throughput, web page loading (both national and international sites), file transfers using fixed-size files, and video streaming quality.

On top of the controlled tests, umlaut layers a crowdsourcing component — the „User Experience Test“ — drawing on billions of background samples collected from real devices over many weeks. This measures broadband coverage, real-world download speeds, latency and time spent connected to 4G and 5G. The results are combined into the umlaut Score on a 0–1000 scale, translated into school-style grades (outstanding, very good, good, and so on). In its European benchmarks, the controlled voice and data tests carry most of the weight, with crowdsourcing contributing roughly a quarter of the final score. Umlaut publishes its full method in its Telecommunications Benchmarking hub and the detailed 2025 Mobile Benchmarking Framework white paper.

Ookla: the crowdsourced speed referee

Ookla takes the opposite approach. Its data comes almost entirely from consumer-initiated Speedtests — the millions of people who open the app or website to check their connection. That produces enormous scale: more than 250 million tests per month, capturing over 1,000 attributes per test across real devices, locations and network conditions.

Each test measures three fundamentals: download speed, upload speed, and latency (ping). For its rankings, Ookla distills these into proprietary metrics. The Speed Score combines download and upload speeds — historically weighted heavily toward download, at roughly a 9:1 ratio — with newer awards also factoring in loaded latency. The Consistency Score measures the share of samples that clear a minimum speed threshold, rewarding networks that are reliably fast rather than occasionally spectacular. For fixed broadband, the Connectivity Score blends raw speed (50%) with video streaming (25%) and web browsing experience (25%). To even qualify for a Speedtest Award, a provider must account for at least 3% of samples in a country and cannot be an MVNO.

The strength here is authenticity and scale: these are the speeds real people experienced on their own phones. The weakness is bias — tests cluster where users already have good signal and choose to run them, so pure crowdsourcing can under-represent dead zones and rural gaps. Ookla documents exactly how the awards are calculated in its Speedtest Awards methodology.

P3: the methodology that started it all

If umlaut’s method feels rigorous, that is because it is the P3 method — inherited wholesale. P3 pioneered the combination of drive testing, walk testing, railway testing and crowdsourcing into a single holistic score. When people still refer to „P3 testing,“ they are almost always describing the same drive-and-walk-plus-crowd approach now branded umlaut (and, most recently, simply „connect“ under Accenture). Treat P3 as the historical name for the benchmark, not a separate living methodology. The lineage is confirmed in the official P3-becomes-umlaut announcement.

Side-by-side comparison

AttributeUmlaut (connect)Ookla (Speedtest)P3 (legacy)
What it isIndependent holistic benchmark, now under AccentureCrowdsourced speed measurement platformOriginal benchmark; rebranded to umlaut in 2019
Primary focusMaximum network capability plus real-world experienceReal-world consumer-experienced speed and reliabilitySame holistic focus as umlaut (its predecessor)
Data collectionDrive tests, walk tests, railway tests + crowdsourcingConsumer-initiated app/web tests (250M+/month)Drive, walk, railway tests + crowdsourcing
ConditionsControlled, scripted, repeatableUncontrolled, organic, user-triggeredControlled, scripted, repeatable
Evaluation criteriaVoice (call success, setup time, POLQA speech quality); Data (download/upload, web, file transfer, video); Crowd (coverage, latency, 4G/5G time)Download speed, upload speed, latency, jitter; consistency vs. thresholds; video & browsing for fixedVoice, data and crowd KPIs — identical framework to umlaut
Scoring modelumlaut Score, 0–1000, mapped to school gradesSpeed Score, Consistency Score, Connectivity ScoreP3 Score (historical, 0–1000 basis)
Coverage of dead zonesStrong — teams actively drive/walk weak areasWeaker — samples cluster where users have signalStrong — same active testing model
Video experienceMeasured inside controlled data tests (streaming quality, start time, stalling)Dedicated Video Score; video weighted 25% in the fixed Connectivity ScoreSame as umlaut — video within data testing
Gaming / real-time responseCaptured via general latency KPIs; no dedicated gaming scoreStrongest — dedicated Game Score using median game latency, jitter, download & upload, measured to real game serversLatency KPIs only (legacy)
AI / agentic response timeNo dedicated metric; latency and consistency serve as proxiesNo dedicated metric yet, but latency-first focus + Accenture’s „agentic, low-latency“ direction lean this wayNo dedicated metric
Best forOperators, regulators, „best network“ certificationConsumers, marketers, fast broad snapshotsHistorical reports and legacy references

What about video, gaming, and AI response times?

This is where the two philosophies diverge most sharply. Ookla has leaned hard into experience-specific scores: a dedicated Video Score for streaming quality and a Game Score built specifically around responsiveness — median game latency and jitter measured to real game servers, blended with download and upload. Because gaming and video are about how quickly and consistently packets arrive, and Ookla’s whole model is latency-and-consistency-first, it is the natural home for these metrics.

Umlaut (and therefore legacy P3) does test video streaming, but as one component inside its controlled data suite rather than as a headline consumer score, and it has no public gaming-specific rating — its edge is breadth and rigor of coverage, not experience-tier granularity.

And AI response time? It is worth being precise here: none of these network benchmarks yet publishes a dedicated „AI responsiveness“ score. Metrics like time-to-first-token belong to LLM benchmarking, not network testing. What the network testers do measure — latency, jitter and consistency — is exactly what determines how snappy an AI assistant or agent feels over a mobile connection. Given that Accenture framed its Ookla acquisition explicitly around „agentic access“ and „low-latency, zero-friction connectivity,“ expect an AI-responsiveness metric to emerge from the Ookla/umlaut camp before long. For now, treat latency and consistency as the honest proxies.

So which one should you trust?

It depends on the question you are asking. If you want to know how good a network can be — including in the places people rarely test — umlaut’s controlled drive-and-walk approach is hard to beat, which is why operators and regulators lean on it. If you want a fast, honest snapshot of what typical users are actually getting today, Ookla’s crowdsourced Speed and Consistency Scores are the most representative of everyday experience. And P3? It is the historical root of the entire controlled-testing school — worth knowing when you stumble across an older report.

The most interesting development is that this distinction may soon blur. With both umlaut and Ookla converging under Accenture, the long-standing tension between controlled benchmarking and crowdsourced measurement could turn into a single combined offering — powerful for operators, but a reason for the rest of us to keep asking exactly how any „best network“ badge was earned.

Sources & further reading

Jan Claude Drasnar
Jan Claude Drasnar
Articles: 14