Conflicting speed scores

PageSpeed Insights says failed. GTmetrix says great. Which result should you trust?

Usually both. They measured different things. Don't average the scores or pick the one you like: check whether they measured the same metric, source, page, device, time window, and test conditions.

Illustrative numbers. Both can be true at the same time.

Neither result wins by default

Neither result automatically invalidates the other. First check whether both reports use the same metric, field or lab source, exact URL or origin-level aggregate (scheme, host, and port), mobile or desktop profile, reporting window, and test conditions.

Field = real visits over time. Lab = one test now.

Use eligible Chrome field evidence to understand observed experience over time. Use repeatable lab tests to try to reproduce the symptom, inspect likely causes, and recheck immediately after a change.

No field data means unknown, not good

If public field data is missing, the result is unknown—not good. Keep a consistent lab baseline, and use first-party real-user monitoring (RUM)—measurements collected from visits to your own site—when the business needs owned visitor evidence.

What each report measured

Two kinds of evidence. Neither replaces the other.

PageSpeed Insights and GTmetrix can each show both kinds. Name the panel before you compare the number.

Field data

What real Chrome visits experienced

p75 is not an average. It is the 75th percentile: at least 75% of included metric experiences were at or below that value—not 75% of unique visitors.

CrUX updates daily over a trailing 28-day period with processing delay. After a change, new visits gradually replace older ones. There is no instant, clean post-change window.

Source: Chrome CrUX API

Lab data

One controlled test, right now

  • Lighthouse, the PSI lab panel, and a GTmetrix test load one URL on a chosen device and network.
  • Great for reproducing a symptom, finding the cause, and rechecking right after a change.
  • It cannot tell you what real visitors experienced, and Lighthouse has no direct lab INP.

Before you trust either result

Six checks before calling the tools contradictory.

Most apparent conflicts compare unlike evidence. If any of these differ, you are not looking at a real contradiction yet.

Same metric

01

An individual LCP, CLS, or INP metric—not a Lighthouse score versus a Core Web Vitals assessment.

Same source type

02

Field compared with field, or lab compared with a comparable lab run.

Same scope

03

Exact URL, origin-level aggregate (scheme, host, and port), or Search Console group of URLs with similar user experience.

Same device

04

Mobile and desktop remain separate.

Compatible time

05

A trailing field window is not the same moment as a lab test run now.

Same lab conditions

06

URL, tool/version, location, device/viewport, network/CPU profile, cache and consent state, and content variant.

Name the panel

Which report are you actually looking at?

“PageSpeed,” “Lighthouse,” and “GTmetrix” are not single data sources.

Field

PageSpeed Insights field panel

Field — Chrome UX Report (CrUX)

URL-level data when available; PSI may show origin-level data instead. Confirm the displayed scope. Mobile or desktop. Trailing 28-day collection period at p75.

Use it to: Judge observed experience for the eligible public Chrome sample over time.

Can't prove: Every visitor's experience, same-day behavior, or the exact page when PSI is showing origin data.

Google PSI documentation
Lab

PageSpeed Insights lab panel

Lab — Lighthouse

One simulated run of the tested URL on the selected mobile or desktop profile.

Use it to: Try to reproduce the symptom and inspect likely causes immediately.

Can't prove: Historical real-user experience, a stable change from one run, or later field recovery.

Google PSI documentation
Field

Search Console Core Web Vitals

Field — CrUX

URLs grouped by similar user experience (or origin groups), with device-specific status and metric issues. Last 28 days at p75.

Use it to: Prioritize affected site sections and inspect representative URLs.

Can't prove: That every exact URL in the group has the same result or that today's change is processed.

Search Console documentation
Lab

Lighthouse in DevTools or CLI

Lab — controlled browser run

The exact tested URL in one run; conditions vary by runner, hardware, network, and page state.

Use it to: Debug repeatable loading and main-thread work under known conditions.

Can't prove: A Core Web Vitals field pass, direct INP, or causality from a single score change.

Lighthouse variability
Lab

GTmetrix performance test

Lab — Lighthouse-powered real-browser test

One tested URL with configurable location, device, connection, and analysis options.

Use it to: Inspect a waterfall, video, Lighthouse diagnostics, and repeatable configured tests.

Can't prove: That its conditions match PSI or that a good lab grade overrules poor field evidence.

GTmetrix features (vendor claim)
Field

GTmetrix CrUX panel

Field — CrUX displayed by GTmetrix

Public CrUX evidence when available; keep its page/origin and device labels attached.

Use it to: Review eligible Chrome field experience without confusing it with the GTmetrix lab grade.

Can't prove: First-party RUM, every visitor's experience, or current behavior outside the CrUX window.

GTmetrix features (vendor claim)

Decision matrix

What are you seeing? Pick it. Get one next step.

Each answer says what the result can support and where the evidence stops. It does not promise a pass, recrawl date, ranking, or traffic change.

You see

Search Console is poor; Lighthouse is good

  1. What it means

    A Search Console URL or origin group had poor CrUX field evidence. A good Lighthouse run did not show the same load-time symptom; for INP, Lighthouse does not provide a direct comparison.

  2. Your one next step

    Open the failing device, metric, and URL group. Test one representative affected URL repeatedly with the same lab setup. For INP, inspect long tasks and TBT only as clues.

  3. Don't conclude

    Do not assume every grouped URL currently fails, or call the field issue fixed from one good run.

Illustrative example — not measured results

WordPress: Lighthouse 96, mobile group LCP 4.2 seconds

Lab

Lighthouse

96

Performance score

Field

Search Console

4.2 s

Mobile group LCP

≠

Illustrative numbers, not a Nimo audit: one page-load score is green while the Search Console group has poor historical LCP. Those are not equivalent outputs.

Next: Open the mobile LCP issue and inspect its group scope. Test a representative URL with repeated matching mobile runs; identify the LCP element and load-delay breakdown before changing the image, cache, or theme.

Limit: An individual URL may differ from the group aggregate. One fast run does not clear the group issue, and the Lighthouse score does not identify the field cause.

Illustrative example — not measured results

Shopify: URL field data missing, origin LCP 2.1 seconds

Field

PSI · this URL

No data

URL-level field

Field

PSI · origin

2.1 s

Origin-level LCP

≠

Illustrative numbers, not a Nimo audit: PSI shows good origin-level LCP context while the tested product page has no URL-level field evidence.

Next: Label 2.1 seconds as origin context, not product-page LCP. Use repeated lab runs on that product URL and inspect hero media, theme work, and app scripts. Use first-party RUM if the exact page needs visitor evidence now.

Limit: A good origin metric is not a pass for the product page or all three Core Web Vitals. Do not remove merchandising apps without an owner and business-value review.

Copyable report comparison worksheetFill this in before choosing a fix. Paste it into a ticket or client handoff.
Report A / panel:
Report A URL:
Report A metric / value / units / status:
Report B / panel:
Report B URL:
Report B metric / value / units / status:
Metric (not overall score):
Source type (field or lab):
Scope (URL, origin, or Search Console group):
Device / viewport:
Observation date / collection period:
Lab runner/version, location, network/CPU profile:
Cache, consent state, and content variant:
Preselected before/after run counts / all values / summary (for example, median):
One next check / owner:
What this evidence cannot prove:

Category errors

Do not compare these numbers directly.

Each pair looks comparable. None of them are.

Lighthouse score (0–100)≠Core Web Vitals assessment

Do not compare a Lighthouse 0–100 performance score with a Core Web Vitals pass/fail assessment. They are different outputs.

Lab TBT≠Field INP

Do not compare TBT with field INP as though they are the same metric. Use TBT and long tasks only to investigate responsiveness work.

Origin-level CrUX≠This exact URL

Do not relabel origin-level CrUX as evidence for one exact URL.

Search Console example URL≠Every URL in the group

Do not treat a Search Console example URL as proof that every URL in its group has the same measurement.

Mobile, one lab setup≠Desktop, another setup

Do not compare mobile and desktop, or lab runs with different locations, network/CPU profiles, viewports, tool versions, cache/consent states, or content variants.

Load-time CLS≠Full-visit CLS

A page-load lab run can miss interactions and layout shifts later in a visit. A clean load-time CLS is not proof of full-session stability.

One lab run today≠A 28-day field window

Do not compare one immediate lab run with a trailing field-data window and call the difference a contradiction.

After you find the failing metric

Metric-specific next checks.

LCP

Largest Contentful Paint

2.5 s4 s

For LCP, identify the largest element and whether server response, discovery delay, resource loading, or render delay dominates before choosing a fix.

LCP guide

CLS

Cumulative Layout Shift

0.10.25

For CLS, identify the elements that moved and the late content, font, image, embed, or UI state that changed their reserved space.

CLS guide

INP

Interaction to Next Paint

200 ms500 ms

For INP, use field INP for the user-experience verdict. Use long tasks, interaction traces, and TBT only as lab clues to the main-thread work worth investigating.

INP guide

For any metric, keep the exact URL, mobile or desktop profile, source scope, and observation date attached to the number.

Verify without fooling yourself

Match the rerun before you call the change better.

  1. 1

    Write down the failure

    Record the failing source, scope, device, metric, and observation window before changing anything.

  2. 2

    Lock the test setup

    Before testing, choose equal before/after run counts (for example, five each) and a summary such as the median. Keep all run values. Match URL, runner/version, location, viewport/device, network/CPU profile, cache/consent state, and content variant.

  3. 3

    Make one bounded fix

    Choose one bounded fix tied to the failing metric and name the owner or handoff.

  4. 4

    Rerun and summarize

    Verify that the intended change is live, then repeat several matching lab runs. Record the run count and a summary method such as the median; do not pick the best run.

  5. 5

    Check the later field window

    Review a later compatible CrUX or Search Console window before describing the field status; report movement without claiming causation.

Source: Lighthouse guidance on repeated runs and variability

Where nimo fits

See field and lab evidence side by side, clearly labeled.

  • Nimo's public audit can combine available CrUX field evidence with PageSpeed Insights diagnostics, preserve unavailable values, and prioritize a first fix.
  • A matching rerun can compare the same page after a change. Signed-in plans can store history and run scheduled audits.
  • Nimo does not provide first-party RUM, automatically rewrite WordPress or Shopify code, or replace the deep waterfall and trace workflows in specialist tools.
  • Supported Cloudflare changes remain limited and require explicit approval; code changes remain recommendations or handoffs.

Check a public page

Nimo labels available CrUX field evidence separately from Lighthouse lab diagnostics returned through PageSpeed Insights.

>

Public URL. No signup.

The fine print

Keep these limits attached to the evidence.

11 limits of each data sourceCrUX, Search Console, Lighthouse, GTmetrix
  • CrUX is aggregated public Chrome field data for eligible visits, not every visitor and not first-party RUM.
  • PageSpeed Insights may show URL-level CrUX, fall back to origin-level CrUX, or show no field data. Preserve that scope.
  • Search Console groups URLs it considers to have a similar user experience, then reports device-specific status and metric issues for each group. A group result does not prove every exact URL has the same status.
  • The Lighthouse performance score is not a Core Web Vital or the Core Web Vitals pass/fail assessment.
  • Lighthouse does not provide a direct lab INP equivalent. TBT and long tasks are diagnostic clues, not replacement INP measurements.
  • GTmetrix can show both Lighthouse-powered test evidence and CrUX field evidence. Name the panel before comparing it with another tool.
  • CrUX updates daily over a trailing 28-day period with processing delay. New visits gradually replace older ones; it is not an instant post-change cohort.
  • p75 means the 75th percentile, not an average: at least 75% of included metric experiences were at or below the reported value, not 75% of unique people.
  • Search Console can use an origin-level group when a narrower group lacks data; check the group scope before assigning a problem to a template.
  • A matching lab rerun can show a directional change under those conditions; it cannot prove later field recovery or causality by itself.
  • Missing public field data is unknown, not a pass, and URL-level CrUX may never become available.
Where this decision guide stopsThis guide is not for every measurement job
  • Session-level or same-day evidence from your own visitors; use first-party RUM for that job.
  • A replacement for deep waterfall, trace, backend, database, or authenticated-flow investigation.
  • Automatic WordPress plugin, theme, Shopify theme, or app changes. Nimo does not rewrite that code for you.
  • An instant field verdict after a change. Public CrUX evidence is delayed and can remain unavailable.

FAQ

Quick answers for conflicting reports.

Why can Search Console fail when Lighthouse passes?

Search Console groups URLs it considers to have a similar user experience, then reports device-specific CrUX field status and metric issues over time. Lighthouse is one controlled test of one URL. Use the group result to prioritize the real-user issue, then use repeatable lab runs to investigate a representative page.

Why can GTmetrix and PageSpeed Insights disagree?

First identify the panel. Both tools can show CrUX field evidence, while their Lighthouse-powered lab tests can use different locations, devices, networks, and settings. Compare matching source types and conditions before deciding that the results conflict.

Does a good Lighthouse score mean the page passes Core Web Vitals?

No. The Lighthouse performance score is a weighted lab score, not the Core Web Vitals field assessment. A good lab run can coexist with poor or unavailable field data.

Can a lab rerun prove a Core Web Vitals fix worked?

Matching repeated lab runs can show a directional change under those test conditions. They do not prove that the change caused the movement or that the later CrUX field window has recovered.

What should I do when PageSpeed Insights has no field data?

Treat the field result as unknown. Check whether origin-level data is available, keep a consistent lab baseline, and use first-party RUM if you need evidence from your own visitors now. Do not assume URL-level CrUX will appear.

Ready to label your own evidence?

Run a public-page audit, then check that URL-level, origin-level, lab, and unavailable evidence remain clearly labeled.

Run the free audit

Need the machine-readable version? Use /field-vs-lab-core-web-vitals.md.