Peak-Season Support at a Global Retailer
A global e-commerce brand is rolling out a Claude-powered customer-support and product-recommendation system. Follow the team from launch through a 10x Black Friday surge, a cost-mix review, and a brand-safety incident — making the architecture, model-routing, and guardrail decisions that keep it fast, cheap, and safe.
The scenario
NorthMarket, a retailer operating in 14 countries, is replacing a rules-based chatbot with a Claude system that handles real-time shopper chat and bulk back-office copy generation. The launch team must right-size models, control cost and latency at extreme peak, and enforce brand-safety guardrails — all before Black Friday.
- Architecture
- A 'router' agent hands off to specialist sub-agents (returns, order-status, product-reco, copywriting); the team assumed multi-agent was mandatory. Real-time chat is synchronous; catalog/marketing copy is bulk back-office.
- Existing stack
- Next.js storefront, existing rules chatbot, product catalog in Postgres, message queue for back-office jobs, Datadog for observability. No LLM in production yet.
- Scale
- ~40K support chats/day at baseline; Black Friday peaks historically hit 10x overnight. Back-office: ~2M product descriptions and marketing snippets regenerated seasonally.
- Regions
- Shoppers in North America, EU, and APAC; EU traffic carries data-residency expectations. Traffic follows the sun with sharp regional spikes.
- Budget
- Cost-sensitive: finance caps the pilot at a fixed monthly spend and wants per-conversation cost tracked. Overruns trigger a mandatory model-mix review.
- SLA
- Real-time chat: first-token < 1s, full response typically < 5s. Bulk copy: no per-item latency target, must finish within a 24-hour regeneration window.
Keep this context in mind — later questions build on it, and the situation evolves as you go.
- P1 · Q1Model right-sizing across two workloadsHow confident are you?
- P2 · Q2Cost levers that respect latencySelect 2How confident are you?
- P4 · Q3Latency for interactive chatHow confident are you?
- P7 · Q4Guardrail failure modeHow confident are you?
The situation changes
Black Friday: traffic 10x overnight
Peak weekend arrives and support chat volume jumps ~10x versus baseline within hours, concentrated in EU morning and US evening windows. Per-conversation cost and p95 latency both start climbing, and finance is watching the daily spend line in real time. The team must hold the chat SLA without breaking the budget.
P1 · Q5Preserving SLA under surgeHow confident are you?- P2 · Q6Bulk work during peakHow confident are you?
- P4 · Q7Latency and cost knobs at peakSelect 2How confident are you?
- P3 · Q8Multi-region routingHow confident are you?
The situation changes
A cost spike forces a model-mix review
Black Friday is over, but the finance dashboard shows per-conversation cost ran far above plan — enough to trigger the mandated model-mix review. Digging in, the team finds two things: a large share of chats were served by Opus, and the prompt-cache hit rate was near zero for much of the weekend. The router's design is now under scrutiny.
P5 · Q9Diagnosing the cost overrunHow confident are you?- P1 · Q10Architecture reassessmentHow confident are you?
- P2 · Q11Structural cost reductionsSelect 2How confident are you?
The situation changes
A brand-safety incident in generated copy
A week later, a customer screenshots a promotional snippet that made an unverifiable health claim about a product — one of the 300K batch-generated snippets that went live. Legal and brand are alarmed. Investigation shows the batch copy pipeline published straight to the storefront without passing the brand-safety guardrail that the interactive chat path uses.
P7 · Q12Root cause of the incidentHow confident are you?- P6 · Q13Hardening the pipeline post-incidentHow confident are you?