Calibre shows performance from three sources at once: Synthetic testing, the Chrome User Experience Report (CrUX), and your own Real User Monitoring (RUM). They rarely report the same number for the same metric, and that is expected. Each one measures a different thing, in a different way, for a different set of page loads.
This page explains why the numbers differ, so you can read all three with confidence.
Three data sources, three jobs#
- Synthetic runs a controlled, repeatable test from a fixed device, network and location. It is the signal you use to catch regressions before release and to diagnose a page in detail.
- CrUX is Google's field dataset of real Chrome users, reported at the 75th percentile, and the same data Google uses as a search ranking signal.
- RUM is field data from your own visitors, collected in near real time across all browsers, and segmentable by page, device, country and more.
Synthetic is one controlled measurement#
A synthetic test runs under fixed conditions: the same device profile, the same throttled network, and the same location every time. That control is the point.
Because nothing changes between runs except your site, synthetic is:
- Reproducible, so you can compare like with like over time.
- Early, ideal for spotting a regression the moment it appears, well before it reaches a single user.
- Diagnostic, running full Lighthouse audits that tell you why a page is slow, not just that it is.
The trade-off: a single test is one page load under one set of conditions, not the spread of devices and networks your real audience uses.
CrUX is Google's field dataset, at the 75th percentile#
CrUX reports what real Chrome users experienced, at the 75th percentile. A few things shape what it can tell you:
- It only includes users who have opted in, and it does not include iOS.
- A page or origin needs enough traffic to be eligible.
- It is public for any eligible origin, so you can benchmark competitors, not just your own sites.
- In Calibre you can follow it as a trend over time, not just a single snapshot.
Google's Core Web Vitals assessment, the part that feeds search ranking, is evaluated over the most recent 28 days. A change you ship today has to work its way through that window before it counts. CrUX is the right source for tracking the search ranking signal, but not where you watch for an immediate regression.
RUM is your real users, in near real time#
RUM collects performance from your own visitors as they browse, across every browser, and updates far faster than CrUX. It is a far deeper data set than CrUX: every qualifying session, rather than a sample of opted-in Chrome users.
That depth is what makes it practical for triage:
- Segment by any dimension: page, device, country, browser and more.
- Attribute to the elements responsible for each metric.
- Find the biggest issues first: sort by sessions to see what affects the most visitors, then segment to find exactly who is affected.
Why the numbers don't line up#
Different populations of page loads#
Each source watches a different set of page loads:
- Synthetic measures one scripted load.
- CrUX measures a sample of opted-in Chrome users.
- RUM measures your sampled visitors across all browsers.
Three different populations produce three different numbers.
Different aggregation windows#
The same metric is summarised over a different window in each source:
- A synthetic result is a single run (Calibre runs several and keeps the most representative).
- Each CrUX data point reflects Google's trailing 28-day collection, reported at the 75th percentile.
- RUM aggregates over the period you select.
Summarise the same metric over different windows and it will not match.
Lab versus field for CLS and INP#
Some Core Web Vitals cannot be measured the same way in the lab and the field. Interaction to Next Paint needs real interactions, so synthetic testing reports Total Blocking Time as its lab proxy instead. Cumulative Layout Shift in the lab only covers the synthetic load window, while in the field it also captures shifts that happen as a real user scrolls and interacts.
Throttling: runtime versus simulated#
Calibre applies network and CPU throttling at runtime during the test. Tools such as PageSpeed Insights estimate the effect of throttling after an unthrottled run. Different throttling approaches produce different timings. See Emulate PageSpeed in Calibre for the detail.
Geography, device and network#
A synthetic test runs from one location on one device profile. Your real audience is spread across countries, devices and connection types, each with its own performance. The wider that spread, the more field data will differ from a single lab run.
How to read them together#
Use each source for what it is best at:
- Synthetic to catch and diagnose regressions before release.
- CrUX to track the search ranking signal.
- RUM to confirm how a change landed for real users.
When a Lighthouse improvement shows up in synthetic, the next question is whether it translated into a better experience in RUM and CrUX. That cross-check is the reason to have all three in one place.