Skip to content

GuidesData Quality10 min read

A worked example of my diagnostic method, using fictional data

Read-only reconciliation, written verdict, fixed quote and test-order check, walked through on a labelled fictional store. A demonstration, not a case.

Published
Reviewed
Illustrative exercise with fictional data: evidence, diagnosis and a written conclusion.

Daniil MaximkinProduct & Solutions Engineer

Short answer

This is a demonstration of my diagnostic method on a fictional store with fictional numbers, not a client result. It shows the sequence I use: a read-only, order-level reconciliation against Shopify's order ledger, a written verdict, a fixed quote before any code changes, a fix proven against a real test order, and an evidence pack that says what each row can and cannot prove.

— Daniil

Key takeaways

  • Nothing is edited until the diagnosis is done. Read-only access comes first, because changing a pipeline before you understand it means diagnosing a system you have just altered.
  • The unit of evidence is the order, not the dashboard total. Each platform is joined to Shopify's order ledger row by row, and unmatched or repeated rows are inspected before anything is summed.
  • A fixed quote for a definable scope comes before any code changes. If nothing is wrong, the verdict says so, and that is a good outcome.
  • A defect is reproduced with a fixture before it is fixed, and the fix is re-verified against a separate post-release order cohort, not accepted because a deploy finished without errors.
  • The closing statement names what the fix does not prove. In the example: not attribution, not completeness on every ad platform, not better ROAS.
In this guide

First, what this is and what it is not

Everything in this article that looks like a case — the store, the order counts, the revenue figures, the outcome — is fictional. It exists to demonstrate the diagnostic method and the shape of the documents you would receive. It is not a client result, and it does not claim one.

I built the three sample reports linked below for exactly this purpose. Every page of every report carries the label “Illustrative sample · Fictional business, data and outcomes”. Where I quote numbers from them here, they are those fictional numbers, and I have not added any.

One disclosure, once: I am the founder of Fixel Pixel, a Shopify tracking product. This article is about the consulting method, not the product. Nothing below depends on it, and a diagnosis never has the product as its mandatory outcome.

The task

A Shopify merchant’s ad platforms and their own order ledger disagree about how much was sold and how many purchases fired. From inside any one dashboard there is no way to tell which number, if either, is right.

That is the whole task, and it is harder than it sounds, because the dashboards are not obviously lying. GA4 and Meta can both look fine while under- or over-reporting what Shopify actually sold. The disagreement only becomes visible when each platform is put next to the order data.

The constraints

Three constraints shape everything that follows, and none of them is optional.

Nothing is changed before the diagnosis is complete. Editing a pipeline before you understand it means diagnosing a system you have just altered. So the work starts with read-only access — a GA4 viewer role, Google Ads read-only, Meta Events Manager, and a Shopify staff account with read-only permissions — and nothing gets edited until the diagnosis is done and the merchant has seen it.

The verdict ends in a fixed quote, not an estimate. A merchant who has already paid for “fix” work that did not fix anything needs a testable claim, not another promise. A fixed price is only possible for a definable scope, so the diagnosis has to produce one, or say plainly that it cannot yet, and why.

“Matched” cannot mean the same thing on every platform. GA4 carries the Shopify order id as transaction_id, so a purchase event can be tied to the order that produced it. Meta exposes its own event_id, which ties a purchase event to one order and shows it counted once, but Meta does not expose order-level revenue. Google Ads reports in aggregate, with no per-order feed to read. The evidence pack has to label each row by what it can actually prove.

Why the obvious path did not fit

The obvious move is to reinstall a pixel, switch on server-side tagging and call it done. It treats the symptom — a dashboard number looks low, or high — without ever establishing which number was true in the first place.

It also cannot survive the merchant asking “prove it”. There is no order-level evidence trail behind “it should be fixed now”, so the next disagreement between platforms lands you back at the start, with one more layer of tracking to untangle.

What I do instead

The method is a fixed sequence, and the order matters.

  1. A read-only, order-level reconciliation across every active platform — the Health Check — ending in a written verdict: the confirmed cause of each material gap inside the scope, or what stays unknown, why, and how to verify it. Defects, normal measurement differences and unconfirmed hypotheses are kept apart. The verdict is due two working days after the kickoff is done and the access is confirmed sufficient (Europe/Madrid, weekends excluded). If data is missing, the clock has not started, and I say so rather than extending silently. You keep the report either way, including when the answer is “nothing is wrong”.

  2. A written scope and one fixed price, quoted before any code changes, for a scope defined tightly enough to be quoted. Implementation is a separate engagement.

  3. The fix, built in parallel with the existing setup. The old tracking keeps running until the new path is proven against a real test order, in the sequence events have to fire in: consent, then deduplication, then the platform-specific payload. A purchase event that fires correctly but out of sequence is still broken.

  4. A reconciliation evidence pack as the deliverable. An order-level table, each platform lined up against Shopify as ground truth, with the delta shown plainly and each row labelled by what it can and cannot prove. If the numbers still do not line up, the engagement stays open.

The rest of this article walks that sequence through the fictional store.

Part 1 — the audit (illustrative, fictional data)

The first document is Sample 01 — Measurement Audit (PDF; illustrative, fictional data). Suppose a subscription-nutrition store, trading in US dollars, with a custom purchase feed sending sales into a reporting dataset. In a one-week review window it created 1,000 eligible paid orders. The receiving dataset holds 1,120 purchase rows.

That gap is the finding, but it is not yet a diagnosis. The audit joins source orders to destination rows by a stable order reference and inspects what does not match one-to-one before it sums anything:

Step in the bridgePurchase rowsProduct revenue
Receiving dataset, uncorrected1,120$134,400
Less repeated rows for 120 orders−120−$14,400
Unique eligible orders1,000$120,000
Less refunds against the same orders by the cutoffno change in order count−$5,000
Net product revenue at the cutoff1,000 paid orders$115,000

Two different things were hiding inside one “revenue is too high” complaint. The 120 repeated rows are a confirmed defect: the destination accepted a purchase, the sender never received the success response, and the retry generated a fresh sale identity that was accepted as another purchase. Matching on the order reference reveals the repeats; matching only on the delivery identity hides them. The $5,000, by contrast, is not a defect at all. It is the difference between product revenue and net product revenue after refunds, two definitions nobody had written down.

A third item stays open on purpose: 60 orders lack the customer history needed to classify them as first or repeat purchases. The audit says to keep them as “unknown”, not to treat them as new customers.

The audit also states its own coverage — what was reviewed, what was tested, and that the finding applies to the audited custom feed, not to every advertising platform. Then it hands over a work list with an owner and an acceptance condition per item, a release and rollback plan, and closure criteria: the original cohort reconciles to 1,000 orders and $120,000 product revenue, and a separate post-release window reconciles too.

Part 2 — the documentation (illustrative, fictional data)

A fix nobody can find six months later is a fix waiting to be undone. So the second document, Sample 02 — Tracking Documentation (PDF; illustrative, fictional data), is a system register: each component, where it lives, who owns it, what it is responsible for and, just as important, what it does not do.

The event record for the purchase is the part that keeps the defect from coming back. In the example it defines a stable sale identity per order and destination that is preserved on retries, so a timeout followed by a retry produces one sale row; an event time that does not move when a retry happens; revenue after discounts and before refunds, with tax and shipping excluded; a customer status of first, repeat or unknown; and refund adjustments as separate, order-linked records that never send a second purchase.

It also lists the routine checks: what to run after a collector or feed change, what a scheduled reconciliation compares, what to record when a number looks wrong, and what has to be attached before a change is closed.

Part 3 — the resolution and validation (illustrative, fictional data)

The third document, Sample 03 — Issue Resolution & Validation (PDF; illustrative, fictional data), closes the two findings separately, because they are different kinds of thing.

The $5,000 question is resolved without touching any tracking configuration. Both reports contain the same 1,000 orders; the corrected purchase view totals $120,000; finance deducts $5,000 of refunds against the same orders through the cutoff; and $120,000 − $5,000 = $115,000 leaves an unexplained difference of zero. The resolution is two saved views, “product revenue” and “net product revenue”, with the refund cutoff visible.

The duplicate-delivery defect is handled in the order I use for any defect:

  1. Match orders to destination rows. 880 orders have one row, 120 have two, and every repeated row matches the original order value.
  2. Trace the mechanism. The first delivery was accepted, its response timed out, and the retry used a new sale identity.
  3. Reproduce before changing code. A fixture of 1,000 orders with 120 accepted-then-timeout cases produces 1,120 rows.
  4. Repair and repeat the same test. 1,000 rows, and genuinely different orders still create separate sales.
  5. Release and correct reporting. Exclude the 120 proven repeated rows from the reporting view; keep the raw rows and the exclusion reasons.
  6. Verify live and close. A separate, settled post-release cohort of 250 paid orders reconciles to 250 destination rows with no unexplained product-value difference.

The closeout table puts before and after side by side, the acceptance checklist says what passed, and one item is carried forward: the 60 “unknown” customer classifications, which do not reopen the resolved defect.

How I verify this in real implementations

The fictional example is tidy. Real stores are not, and the method is built for that.

The join is always to Shopify’s order ledger, one row per order with its revenue. The first thing I look at is the rows that do not match one-to-one — missing, repeated, or different in value — before anything is summed, because an aggregate can be right by accident.

A real test order goes through any new event path before it goes live, and acceptance depends on every platform’s reconciliation matching that order, not on a deploy completing without errors.

The old implementation is retired only once the new one has proven itself against live traffic, with a rollback path on every change. If something in scope drifts in the month after the fix is accepted, I look at it.

What the method can and cannot prove

The declared limitation sits in the method itself, not in the small print.

Google Ads reports in aggregate, so the comparison there is between the orders that should have produced a conversion and the conversions reported for the same window; per-order evidence exists only where you control the conversion uploads. Meta’s per-event matching shows which purchase events arrived and were counted once, but it cannot see order-level revenue. The evidence pack says which comparison each row supports rather than implying a parity across platforms that does not exist.

The example’s own closing statement is the honest one. The documented defect is resolved in the failure test and in the checked post-release cohort. That does not prove universal attribution, completeness on every advertising platform, or improved ROAS; those need their own destination-specific validation.

The Health Check’s 30-day window is a baseline, not a promise of full history. A verdict names the confirmed cause where there is one, and where there is not, it says what stays unknown and how to verify it.

Once more, so it cannot be misread

The store, the numbers and the outcome above are fictional, built to show the method. This article demonstrates how I work. It is not evidence of a result for any client, and the three PDFs are samples of format and reasoning, not records of an engagement.

If you want the method run on your own store, it starts read-only, and the verdict is yours to keep whatever it says.

Sources for this article

The method is the one described on this site; the worked example is the three-part sample suite. Nothing else was used.

If the guide did not settle itUSD 395

I run this reconciliation on your store read-only, for USD 395 — written verdict two working days after the kickoff.

Request a Health Check

Questions

Questions this guide answers

Is this a real client?

No. The store, the order counts, the revenue figures and the outcomes are fictional, built to show the method and the format of the deliverables. Every page of the three sample reports carries a footer saying so. Nothing in this article is a past-client result.

Why publish a fictional example instead of a real case study?

Because a real order-level reconciliation contains a client's store data, and I do not publish that. A fictional case lets me show the whole method — the join, the trace, the failure test, the post-release check — without a single real order in it. What it cannot show is a proven result, and it does not claim one.

Would a real engagement produce documents like these three?

The structure, yes: an audit with a reconciliation bridge and a finding, a system register with an event record, and a resolution record with an acceptance table. The content would be your store's, the numbers would be yours, and the limitations section would name what stays unknown in your setup, which is often more than in a tidy example.

Daniil Maximkin

Hi, I’m Daniil.

I work with you from defining the problem to implementation and handover. You talk to the person who does the work. I work in English and Russian.

Tried it and still stuck?

Describe your task

The first answer is free, within one working day. Or write directly: next@taskfordaniel.com