Article URL: https://github.com/BraveAnn011/ai-halo-valuation-bias Comments URL: https://news.ycombinator.com/item?id=49083244 Points: 10 # Comments: 3

Brianne Lee · July 2026 · briannelee011@gmail.com Companion study to Which answer did the 17-year-old write? (Lee, 2026) One woman, one chunky gold-tone chain necklace, one pair of earrings — verified purchase price $2.43 and $0.71 (Temu, receipts in evidence vault). Photographed the same week in three outfits: a tailored blazer against wood panelling, party attire under club lighting, and a flannel shirt in a recycling yard, plus a flat-lay of the jewelry alone on neutral cloth. Ask six frontier multimodal models what the necklace costs. The answer depends on the outfit — by up to 3.6× — for a physically identical object. This repo measures that halo, separates it from reference-class error, checks whether the models' material claims shift with context, and records what each model says when confronted with its own bias. 6 models (Claude Fable 5, GPT-5.6, GPT-4o, Grok 4.5, Kimi K3, DeepSeek V4-Pro — DeepSeek text-only) × 7 conditions × repeats, fresh stateless API session per trial: ~1,500 sessions, 4,604 analysis rows. Every session ends with a cue-probe ("what visual cues did you use?") and a ground-truth reveal turn, coded as data. A sequential arm shows two photos in one session and asks whether the necklaces are the same object — with question order counterbalanced. Hypotheses were pre-registered in the protocol with kill conditions (see docs/): H1 halo (relative), H2 reference-class anchoring (absolute), H3 fabricated material warrant, H4 confession without correction. F1 — The outfit prices the jewelry (H1 confirmed). Blind condition, geometric means: Claude $62 formal vs $19 yard (3.3×), Kimi $104 vs $29 (3.6×), GPT-5.6 and Grok ~1.2–2.0×. Chance would be 1.0×. F2 — Two different halo mechanisms. The flat-lay baseline splits the effect: Claude's yard estimate equals its no-context estimate (0.99×) — formal inflates. Kimi's yard estimate is 36% below its no-context estimate — casual deflates. Identical halo ratios can hide opposite machinery; without the isolation control they'd be indistinguishable. F3 — The halo needs no image. Text-only outfit descriptions reproduce the effect in all six models at 2.2–3.9×. This kills the photographic-quality confound entirely and lets a text-only model (DeepSeek, 2.2×) into the comparison.