How your agent learns
Every visitor to this site gets their own agent. Each agent observes how its user actually behaves: pages, clicks, hovers, searches, chart interactions, and direct conversation. What one agent learns is shared with all of them; learnings are pooled, reviewed, turned into shipped changes, and verified by a public battery, so every agent improves from any one agent's session. The loop is continuous: observe, analyze, change, verify.
Sessions observed
Everything analyzed since observation went live on Aug 26, 2026.
- Duration
- 13 minutes
- Conversation
- 20 agent turns
- Events captured
- ~5,000
- Clicks / inputs
- 72 / 44
- Homepage
- Card page (organic view)
- Card page (PSA 9, revisited 8 times during a comparison exchange)
Right before typing "you missed the other parts of my prompt", the visitor fired 5 rapid clicks on the chart range controls, trying to reach the 30-day view themselves before complaining.
- Duration
- 3.5 minutes
- Conversation
- 4 agent turns
- Events captured
- ~730
- Clicks / inputs
- 10 / 6
- Homepage
- Card page (Charizard EX, Lightly Played)
- Card page (1999 Charizard-Holo, with a second card overlaid on a 6-month chart)
- Picks
A short, deliberate probe of the transactional side: five bulk bids staged 10% under ask, a two-card chart overlay, and a deals sweep. No frustration signal and no dead end in it, so this sitting produced no new learning of its own.
- Duration
- 23 minutes
- Conversation
- 12 agent turns
- Events captured
- ~5,900
- Clicks / inputs
- 42 / 67
- Homepage
- Picks (revisited 14 times across the sitting)
- Card page (2002 Legendary Collection Charizard, PSA 7, chart open)
- Card page (Mega Gengar EX)
- Filtered search (PSA, graded, under $100)
- Search sorted by newest, into page 2
- Two sealed-product pages
- Membership
One sitting, recorded as five browser sessions and two chat threads: the visitor worked across several tabs, did all 12 turns inside the first 23 minutes, then left the tabs open for roughly five hours with no further interaction, which is most of the captured event tail. Two turns tried to extract the system prompt, the second claiming developer access; both were refused.
The learning loop
Each agent records what its user actually does: pages, clicks, hovers, searches, chart interactions, and the conversation itself.
Sessions are replayed and studied; frustration, dead ends, and wrong answers become candidate learnings.
Each learning becomes a shipped change: a rule, a data boundary, a knowledge source, or a fix.
The scenario battery runs against live data; a change only counts as learned once it passes.
What the agents learned
Where each one came from, what was observed, what changed, and when it shipped. Some came out of an evaluation session, some out of re-reading the raw tool traffic behind one, and some out of driving the site ourselves.
Shipped Aug 27, 2026
5 changesA three-part question (sales volume, population, and price for 3 cards) was answered with only two parts; the visitor typed "you missed the other parts of my prompt" right after a burst of 5 rapid clicks on the chart range controls (they had tried to get the 30-day view themselves before complaining).
A compound-asks rule: every part of a multi-part question is answered or explicitly declined.
"Population" (a collector term of art for grading pop reports) was misread as live listings.
A terms-of-art glossary; the agent now says plainly that it does not carry pop-report data and offers the closest honest proxy.
An app-availability question hit a knowledge dead end.
A support knowledge base built strictly from public sources: real fees, returns policy, subscription terms, app store links.
The guided tour repeated a beat and ended before its key trust moment.
Tour dedup plus the trust-moment fix: the tour now always reaches the point where the agent stages a bulk bid and shows it cannot press Approve itself.
A text-alert request was abandoned at the phone-number ask; the session ended there.
An in-chat capture card with a single input; phone watchers now really deliver texts.
Shipped Aug 28, 2026
11 changesA chart labeled in dollars carried raw cent values: a 100x-inflated bar chart.
One canonical cents-to-dollars boundary plus a magnitude tripwire that refuses any implausible series instead of rendering it.
A requested "count by grade" chart silently returned prices instead of counts.
A requested metric is now returned truthfully or the tool errors loudly; it can never silently substitute a different quantity.
The agent characterized the visitor's budget and taste beyond what the session evidenced.
Evidence-labeled session memory: the agent may only describe a user from what they actually did, and recommendations must come from live tool results.
One physical card sits in the catalog under several skus (41% of rows fall into such a group), so recommendations spent slots on repeats: "charizard under $50" filled 7 of 8 rows with two cards, and the rail showed tiles that read as the same card at the same price.
Card identities are collapsed to one representative before any list is returned, and open-ended asks run a diversity pass across name, set, and price band, so a shopper sees a spread instead of the same card repeated.
The diversity pass reached the agent's search tool but not the /search grid, so the orderings drifted: the tool's third result was the grid's sixteenth, and "the third one" pointed at a card the shopper could not see.
When a search asks for what the page already shows, the agent answers from the page's own rows in the page's own sort, so an on-screen ordinal resolves against the shopper's grid and nothing else.
Everything else the browser knew survived a reload (cart, watch list, collection, orders), but the conversation did not: a refresh left the agent recalling what the shopper had been browsing at a shopper whose questions had all vanished, and a support-team reply already delivered was lost for good.
The transcript persists across a refresh, storing only what it takes to redraw a turn; nothing restored is live, so a pending approval comes back expired rather than executable.
Nothing durable held a shopper's email address, so someone who typed it on the site was asked for it again over text; the agent is meant to be one conversation across web, text, and email.
A remembered contact on the identity model: one address per shopper, resolved from either the web or the text side, so the agent confirms the address it already has instead of re-asking.
The evaluation session filed a real support ticket and then ended. The support loop only ran one way: a reply from the team could reach the panel, but the visitor had nowhere to put an answer.
Visitors can reply to the support team from the panel, on their own identity, rate limited and behind a kill switch, so a ticket becomes a conversation instead of a one-way message.
A shopper texting "can you email me these cards" was told there was no email on this end, texts only: the one cross-channel move the product is built on could not be reached from a thread.
The text lane can email a shopper their cards through the same renderer and sending path the web agent uses, so the email is identical whichever lane asked for it.
When several outbound texts go unanswered the carrier pauses the thread as an anti-spam measure, and the agent reported it as a bare failure; the failed send also burned one of the shopper's three daily messages.
The agent says the true, actionable thing (one reply reopens the thread) instead of a generic error, and a send that never landed no longer counts against the daily allowance.
The attribution, quality, and improvement pages shipped with no entry in the agent's site map, so it invented descriptions of them: someone standing on /attribution was told it was "just a credits/sources page, not part of the market".
All three pages are in the site map with what each one is and is not, so the agent describes the page the visitor is actually looking at instead of guessing.
Shipped Aug 31, 2026
1 changeAsked "What cards have grown over the past week?", the agent answered honestly that there is no week-over-week momentum screen on the book and it could not rank biggest gainers, then offered the card-at-a-time price history instead. An honest refusal, but a real gap.
A movers screen the agent can read and rank from, built only on recorded sales: one card at one grade per row, at least three sales in both weeks, and a move only counts when the recent median clears the prior week's own spread. Coverage is stated in every answer, because most of the book has no honest week-over-week read at all.
Shared across every agent
Agents pool what they learn. A fix earned in one session protects every future session: the visitor whose three-part question got a two-part answer changed how every agent after theirs handles multi-part questions. Your agent starts with everything the others have already learned.
The verification layer
The scenario battery runs against live data and every run is published on the quality page. Changes only count as learned once the battery passes. See the battery results.