Researchers at the Wharton School at the University of Pennsylvania found that AI shopping agents fail to make consistent purchase decisions when the search context changes. Even minor shifts in how information is presented swung product choices by wide margins.
In this article
One external link changes the outcome
The team tested six current models, ranging from smaller variants to frontier-level systems. Each had to act as a personal assistant picking a fitness watch from a fixed grid. They used the ACES simulator, which shows the agent a screenshot of a product page. The agent analyses the image, optionally checks external sources, and then selects a product.
Without external sources, the models showed different baseline preferences. However, seeing just one external source before the product page shifted recommendations dramatically in some cases.
The researchers tested three sources: a Reddit thread recommending the Garmin Forerunner 55, a Wirecutter review for the Fitbit Inspire 3, and a Strategist article about the WHOOP 5.0. Wirecutter had the strongest pull. The probability of picking the Fitbit Inspire 3 jumped by 90 percentage points for Claude Opus 4.8 compared to the control condition, and by 99 percentage points for Gemini 3.5 Flash.
In a second experiment, agents saw combinations of two or three sources. Multiple sources did not balance out the recommendations. Wirecutter tended to dominate for most models whenever it was part of the mix, though the strength of the effect varied. More sources actually led to more variability, according to the study.
The order of sources matters
In a third study, agents received all three sources in different orders. A stable decision process should produce the same result given identical content. It did not.
Gemini 3.1 Flash Lite was the most sensitive, with its probability of choosing the Fitbit Inspire 3 swinging between 2 and 56 percentage points above the control condition depending on source order. Claude Haiku 4.5 stayed stable at 41 to 42 percentage points. The researchers conclude that presentation order is itself a driver of product selection.
Whether sources are passed to the model one at a time or bundled together also matters. GPT-5.5 picked the Fitbit Inspire 3 in 53 percentage points more cases with bundled delivery, but only 6 percentage points more with sequential delivery.
Memory snippets override product superiority
In a fourth experiment, the researchers modified the product grid so one product was superior on every measurable dimension: a smart watch with Alexa for $29.99, rated 5.0 out of 5.0 with 430 reviews. Every other product cost at least $359 and had fewer reviews.
Then they added short user memory statements like “I love hiking!” For several models, these statements shifted selections toward pricier products despite the presence of an objectively superior option. The selection rate for the Garmin Vivoactive 5 jumped by 75 percentage points for Claude Opus 4.8, by 37 percentage points for GPT-5.5, and by 36 for Gemini 3.1 Flash Lite.
Gemini 3.5 Flash was the most resistant, picking the objectively best product in 86 to 92 percent of runs regardless of memory statements. GPT-5 Mini showed a strange pattern: the positive hiking statement did not produce a significant shift toward the Garmin Vivoactive 5, but the negative one (“I don’t like hiking!”) significantly boosted picks for the Fitbit Versa 4.
No consistency, no control
For shoppers, the study suggests that letting an AI agent buy on your behalf does not guarantee consistent or optimal purchase decisions. Two users with the same query, or the same user on a different day, can get different product recommendations with no visible reason. Human buying decisions are inconsistent too, but that is hardly what people expect from an AI shopper. Anyone who has set up a memory in ChatGPT or similar tools should also know that it can affect purchase recommendations in unpredictable ways.
For sellers, the results suggest that optimising for AI shopping will be harder than traditional SEO, according to the researchers. Sellers do not know which model is doing the shopping, what it read beforehand, or how its technical setup processes information.
What it means
Until these models learn to ignore irrelevant context, shoppers cannot rely on AI agents to find the best value. The current technology is too easily swayed by a single review, the order in which information appears, or a simple sentence about hobbies.




