A model can only get better than the humans it copies if something scores its answers. Retail already keeps that score.
Both ask the same question of one customer at a time — would you accept this? — and both are graded against offers real people did or did not take.
The case studies are what the model does. These are how you call it.
Each one is a real study with a real topline the model was never trained on — the questionnaire, what it was shown, what it was kept from seeing, and how every figure was worked out.
Read the three studiesEggAI ships as libraries you call from your own systems. Each one points the same behaviour model at a different decision, and each has a published study behind it.
Two of the three run on your transaction history. The third needs nothing at all.
Each SDK rests on one of these. Every paper gives the data, the split, the scoring and the result — including where a purpose-built system beats us.
A per-customer search over discount depth and minimum spend, applied to a completed voucher campaign.
Ranking quality on a held-out voucher benchmark, and what it implies for a team that cannot call everyone.
A 27-question health-and-beauty study, re-run against a simulated panel and marked on distribution match.
Why a behaviour model beats a bigger general one at this, in three moves.
An egg drawn twice — amber behind, emerald in front — with one trace running through it to a node. The offset is the echo of a customer; the trace is what we read.
Two brand colours, read as a spectrum. Everything else is ground and ink.
Forty-five minutes on your category, with the panel answering live. No slides unless you ask for them.
Both read one customer’s own order history and answer a single question in first person — would you accept this? One walks the terms down until the answer turns to yes. The other keeps the names where the answer is already yes and drops the rest.
A discount that gets redeemed tells you nothing about the discount you could have got away with. The model answers the counterfactual one customer at a time: at these terms, would this person still redeem?
So the search runs downward. Start at the smallest discount and the hardest minimum spend, read the constraint the twin actually objected to, move only that lever, and stop the moment the answer turns.
The same engine, pointed at a list instead of at an offer. Score every name for predicted acceptance, rank them, and cut the tail — so an outbound team works the top of a list instead of all of it.
The useful part is how hard it is to please. Asked about 12,480 real offers it said yes to 120 — fewer than 1 in 100. Everything else came back as not worth the contact. A short list someone can defend cutting is worth more to a rep than a long one nobody finishes.
Both use cases rest on the model preferring some offers to others, so we tried to trip it up. We slipped a meaningless sentence into the question — a line about Thailand having 77 provinces — and the score moved a little. Then we personalised the actual offer, and the score moved about the same amount, while the number of people who said yes did not budge: 120 before, 120 after.
So the honest reading is that this is a sorter, not a persuader. It orders a list well enough to cut it, and it finds the point where someone stops saying yes. It has not shown that changing an offer causes a person to accept. Proving that means holding back a random group and deliberately not sending to them, and we have not done it.
Everything on this page is a prediction, marked against what people really did under offers somebody else had already chosen. Where a number would need a test we have not run, it is not on the page.
A per-customer search over discount depth and minimum spend, applied to a completed voucher campaign.
Discounts are set for segments, so most customers receive more than they needed to convert. We ask a behaviour model, customer by customer, how far an offer can be reduced before that customer stops accepting it, and re-price to the point just above the refusal. Applied to 939 re-priceable offers from a completed campaign, the procedure removes between 6% and 15% of the discount committed to them while holding predicted acceptance fixed.
An offer that is redeemed tells you the terms were sufficient. It does not tell you whether they were necessary. The counterfactual — would this customer have accepted less? — is unobserved for every redemption in a campaign, because only one set of terms was ever sent.
A behaviour model trained on real purchase histories can be asked that counterfactual directly, one customer at a time, at terms that were never offered.
For each customer the search begins at the least generous plausible offer — smallest discount, highest minimum spend — and asks the twin, in the first person, whether it would redeem. The reply carries a likelihood, a decision, and the constraint it objected to.
Only the lever named in that objection is moved, and the offer is re-scored. The search runs up to eight rounds and stops at the first acceptance, which is that customer’s cheapest accepted terms. Where nothing inside the budget is accepted, the procedure returns no offer rather than manufacturing one.
The published range is 6–15%: the floor is the band corroborated by what comparable customers actually did, and the upper figure is set well inside what the model itself reaches. Sales are unchanged in every band, since the procedure holds acceptance and moves only the discount.
Measured against the campaign’s full discount bill rather than the re-priceable subset, the same bands are 1.7% and 16%. The subset carries 27% of the bill.
The binding constraint is reach, not pricing. The 939 re-priceable offers hold 27% of the campaign’s discount; the rest sits on offers the model either did not identify as redemptions or could not price lower. Improving recall on high-value offers moves the ceiling further than improving the search.
Every figure ranks or matches behaviour recorded under the campaign that was actually run. Establishing that cheaper terms cause the same purchase requires a randomised holdout cell, which is the natural next step for a live deployment.
Ranking quality on a held-out voucher benchmark, and what it implies for a team that cannot call everyone.
Outbound teams hold more leads than they have capacity to contact, so the operative decision is which subset to work. We evaluate a behaviour model as a ranker on 12,480 held-out voucher offers with known outcomes. Presented with one person who accepted and one who did not, the model orders the pair correctly 76 times in 100 against a chance rate of 50, and at its own acceptance threshold it retains under 1% of the list.
A list is worked from the top. If its order is arbitrary, the acceptance rate of the first hour equals the acceptance rate of the list, and on this benchmark that is 15 in 100 — the remaining 85 contacts reach someone who was never going to accept.
Ranking is therefore the whole product: not whether a contact will accept, but whether they will accept sooner than the next one.
The model was trained on a different company, in a different country, answering a different question. Each row carries a shopper’s profile and their activity before the voucher was collected; activity after collection is excluded, since it is unavailable at decision time.
Each row is scored by reading the probability the model assigns to acceptance directly, which is the readout the published benchmark uses. Ranking quality is the probability that a randomly drawn accepter is placed above a randomly drawn non-accepter.
A second readout — asking the model to reason in the first person, then judging the written answer — was run on the same 12,480 rows. It is the readout used wherever a closed frontier model appears in a comparison, since such a model will not expose a probability.
On the judged-answer readout, where a frontier model can be included, the same comparison is 69 in 100 for the behaviour model against 64 for a frontier model and 60 for the same engine untrained — with no examples from the target platform at any point.
The model is also markedly selective. At its own acceptance threshold it returns yes on 120 of the 12,480 offers, under 1 in 100, and its mean predicted acceptance across the set is 0.155 against a true rate of 0.153.
A purpose-built system and a conventional statistical model both reach 79 on this benchmark, three points above the behaviour model, which suggests the remaining gap is a property of the available features rather than of the approach.
The split is disjoint by session rather than by customer, so the result characterises ordering among contacts a business already has history for. Scoring a contact with no history is a separate problem: on a cross-market test of 414 consumers with no identity overlap with training, the model reached 71.7% on purchase decisions against a frontier model’s 70.6%.
A 27-question health-and-beauty study, re-run against a simulated panel and marked on distribution match.
Survey research is slow and expensive to field, and a simulated panel is only useful if it reproduces the spread of a real one rather than a single plausible answer. We took a health-and-beauty study already fielded to 1,013 people and put the same 27 questions to a 150-respondent simulated panel. Across all 27 questions the simulated panel places 64% of its answer mass where the real respondents placed theirs, against 52% for a general-purpose frontier model and 47% for the same engine without behavioural training.
A model asked a survey question will produce a plausible answer. A panel produces a distribution — a spread of answers across a population, with minorities and fence-sitters in proportion. Reproducing the first is easy and commercially useless; the second is the thing a research buyer is paying for.
The evaluation therefore marks distributions against distributions, question by question, rather than scoring any individual answer as right or wrong.
Every model receives the same questionnaire in the same format and is marked by the same analyser. For each question, the simulated panel’s answer distribution is compared with the real topline; the reported figure is the share of answer mass that lands in the same place, averaged over the 27 questions.
Elicitation is held constant across models, since the mode in which a panel is asked to answer materially changes the spread it produces.
On the brand-loyalty question in the body-wash category, where 188 real buyers answered, the real study found 38.0% keep to a single brand. The simulated panel returned 41.3%, the untrained engine 10.0% and a frontier model 8.7%.
The frontier model’s characteristic failure is visible on the spending-trend question, where it placed 90.7% of its respondents on one middle option. Ten percent of real respondents chose it.
The gap between the untrained engine at 47% and the trained panel at 64% is attributable to behavioural training rather than to model scale, since both share the same base. The frontier model sits between them at 52% while collapsing on the questions where a population is most dispersed.
Distribution match is an aggregate property. Respondent-to-respondent variation within the simulated panel remains narrower than in a fielded one, which places the instrument at questionnaire piloting, concept screening and segment sizing rather than at replacing fieldwork.