पाठशाला Pathshala · वृद्धि Vṛddhi, Growth · Lesson 22 · Build
Attribution: what actually drove the sale
Every ad platform marks its own homework. Build attribution that follows the customer across phones, shops and WhatsApp chats, and test what the dashboards claim against a group that never saw the ads.
Pathshala, The Founder Library · 11 October 2026 · 7 min read

Add up the conversions that Meta, Google, the affiliate network and the influencer agency each claim for last month and the total is often larger than the number of orders the company actually received. Each platform is reporting honestly by its own rules. None of them is answering the founder’s question, which is what the money caused.
Attribution is the work of answering that question well enough to move budget. In India it has three complications that most global guides ignore. Customers switch devices, often between a personal and a family phone. A large share of journeys pass through a WhatsApp chat, where the trail goes cold. And many sales close offline, in a shop, on a call or at a demo, days after the click. This lesson sets up attribution that handles all three, and then adds the step that most companies skip: testing what the dashboards claim against people who never saw the ads.
Why platform dashboards overstate
A platform sees a person who saw or clicked an ad and later bought, and it counts the sale. It cannot see whether the person would have bought anyway. For a brand with any demand of its own, many would have. The cleanest evidence comes from eBay. Tom Blake, Chris Nosko and Steven Tadelis ran large field experiments on eBay’s paid search, published as Consumer Heterogeneity and Paid Search Effectiveness. Ads on brand keywords, people searching for eBay by name, had no measurable short-term benefit. Ads on other keywords did influence new and infrequent users, but frequent users, whose purchases were not changed by the ads, accounted for most of the spending, so the average return on those ads was negative. Returns measured by experiment were much smaller than the usual non-experimental estimates.
The general lesson holds for every channel that targets people already likely to buy: retargeting, brand search, lookalikes of existing customers, creators whose audience already follows the brand. Each dashboard will look excellent. Each dashboard is also counting customers the company would have had. Measuring the truth is expensive, as the twenty-five large experiments in Randall Lewis and Justin Rao’s The Unfavorable Economics of Measuring the Returns to Advertising show, but a founder does not need precision to the rupee. A founder needs to know whether a channel’s true CAC is near its reported CAC or three times it.
One customer, one identity: the phone number
Cross-device journeys break cookie-based attribution because the click happens on one device and the purchase on another. In India the fix is unusually simple: almost every customer gives a mobile number at some point, at sign-up, on WhatsApp, for delivery or for an OTP. Make it the key of the customer table. Every touch the company can record, a tagged link, a chat, a form, a call, a shop visit, is written against that number with a timestamp and a source. Hash it before sending it to any ad platform, and handle it under your privacy notice and the [DPDP Act lesson](/library/dpdp-act-what-it-requires-of-your-product).
Then record the source at first contact in your own system, not only in the analytics tool. When a number first appears, store how it arrived: the UTM parameters of the landing link, the keyword of the WhatsApp message, the creator’s code, the referrer’s number, or the answer to a question. This field, first touch, never changes. A second field, last touch before purchase, is updated on each order. Two fields in your own database will answer most attribution questions better than any dashboard, because they cover every channel by one rule.
WhatsApp and offline journeys
WhatsApp is where most Indian attribution goes dark: a customer clicks an ad, taps through to a chat, talks to a salesperson for three days and pays by UPI link. Close the gap at the first message. Meta’s ads that click to WhatsApp open a chat directly from Facebook and Instagram, and Meta lists its pixel, Conversions API and offline conversions as the ways to measure results beyond the conversation. Whatever the route, prefill the first message with a campaign keyword, have the CRM read it, and attach it to the number. When the order is paid, send the conversion back to the platform with the hashed number so its optimisation learns from real sales.

For sales that close in a shop or on a call, ask. A single question at billing, how did you first hear about us, with six fixed options and the staff trained to tick one, produces data that is noisy but unbiased by any platform. For business software, add the question to the demo booking form and check it on the first call. Upload offline sales against phone numbers to the ad platforms weekly. Self-reported source undercounts channels people do not remember, such as a display ad, and overcounts the ones they do, such as a friend. Read it alongside the tagged data, never alone.
Models: last click as a floor, not a truth
Attribution models divide credit among touches. Google Analytics now offers three: data-driven attribution, paid and organic last click, and Google paid channels last click; first click, linear, time decay and position-based models were removed in November 2023. Data-driven attribution uses machine learning over converting and non-converting paths and is specific to each advertiser. It is useful, and it is still built from the touches the tool can see, which excludes most WhatsApp and offline steps.
Use a simple rule. Last-click and first-touch numbers from your own database are a floor and a lens: they tell you which channels start journeys and which close them. They cannot tell you what a channel caused. For that you need an experiment. Google’s Conversion Lift describes the method plainly: split the audience into a group that sees the ads and a group that does not, and the difference in conversions between them is the lift. Its geography-based version supports offline data, which matters in India.
A dashboard tells you who saw the ad and then bought. Only a holdout tells you who bought because of it.
Running a holdout without a data team
The simplest holdout is geographic. Pick pairs of similar cities or districts, matched on last quarter’s orders and growth; Pune and Ahmedabad, say, or two sets of pincodes within Delhi NCR. Run the campaign in one of each pair and switch it off in the other for four weeks. Compare orders, measured in your own database by delivery pincode, not in the platform. The difference, scaled to the size of the exposed regions, is the campaign’s incremental effect. Run it on the channel you most suspect: retargeting and brand search are the usual places where the platform’s CAC and the incremental CAC are furthest apart.
Set the figure to your own numbers. At the defaults, the platform reported two thousand conversions on ₹10 lakh, a CAC of ₹500. The holdout shows eight hundred were caused by the ads, an incremental CAC of ₹1,250. If your [payback](/library/cac-ltv-and-payback-the-three-numbers) works at ₹500 and fails at ₹1,250, the channel is losing money while its dashboard looks fine. Sometimes the test runs the other way and finds a channel that the dashboards undercount, often video, creators or radio, because their buyers arrive later and by other routes.
Three rules keep a holdout honest. Run it long enough to cover at least one full purchase cycle, and longer if customers usually buy a week or two after first seeing the brand; a four-week test of a product with a six-week consideration period measures only the impatient. Change nothing else in the test or control regions during the test: no new offer, no local event, no creator campaign that spills across both. And avoid festival weeks and sale periods, when baseline demand swings so much that a real effect disappears in the noise. Platform tools such as Google’s Conversion Lift run user-level versions of the same experiment inside the ad system; Google notes that it is not available to every account and needs a conversation with an account representative. A geographic test needs only your own order data and the discipline to leave one region alone.
The quarterly attribution check
Monthly, from your own database: first-touch and last-touch orders by channel, the share of orders with no known source, and each channel’s CAC on both. If the unknown share is above a fifth, fix the tagging before reading anything else. Quarterly, run one holdout on the channel with the largest spend or the largest gap between platform and database counts, and write down its incremental CAC beside its reported CAC. Keep a running ratio of the two for each channel and apply it to the dashboards between tests. Move budget only on incremental CAC. Once a year, switch off the channel you are most confident in for two weeks in a test region; it is the cheapest way to learn whether that confidence is earned.
Platform features and model options change often; the descriptions above were checked on 11 October 2026. The rupee figures are illustrations.
Sources
- Tom Blake, Chris Nosko and Steven Tadelis, Consumer Heterogeneity and Paid Search Effectiveness: A Large Scale Field Experiment, NBER Working Paper 20171, May 2014
- Randall A. Lewis and Justin M. Rao, The Unfavorable Economics of Measuring the Returns to Advertising, Quarterly Journal of Economics 130(4), 2015
- Google Analytics Help, Get started with attribution (models available; models removed November 2023)
- Google Ads Help, About Conversion Lift
- WhatsApp Business, Ads that click to WhatsApp