Following the Preference, Missing the Optimum: Compliance Without Optimization in AI Housing Recommendation
Source note: Preprints on societal impact.
arXiv:2609.10856v1 Announce Type: new Abstract: Large language models are becoming the first point of contact for consumer search in domains where the stakes are material and the law is explicit. Existing audits show that models steer housing seekers by perceived identity, but none can say what a user loses when a recommender overlooks a suitable option, for want of an enumerated inventory to score omissions against. We audit AI housing recommendation against a verifiable ground truth. For each of 150 synthetic renter scenarios in New York City we build a pool of 120 real listings with known rent, bedrooms and GTFS-computed transit commute, compute the exact set satisfying the renter's stated constraints, and derive its Pareto frontier. The primary outcome assumes no utility function: a recommendation is strictly dominated if the same pool holds a listing cheaper, faster to commute from and no smaller in bedrooms. Across 9,945 calls to three models from two vendors, compliance is near-perfect (1.8% violation against a 66.6% random floor), yet 39.0% of recommendations are strictly dominated, and the dominating listing is a median 900 USD/month cheaper and 3.5 minutes closer. A within-scenario manipulation separates two capabilities usually conflated: changing one sentence moves median recommended rent by 646 USD/month in the correct direction, so preferences are honored, yet recommendations still sit 606 USD/month above the five cheapest qualifying listings on the same screen, and an unambiguous lexicographic instruction gives no improvement under equivalence testing against
Every model that read this
| Model | Provider | Stage | Score | Conf. | Latency | Prompt | When |
|---|---|---|---|---|---|---|---|
| Llama 3.3 70B | Meta | analysis | -75 | 90% | 4017ms | v1.0.0 / m1.0.1 | 2026-09-11 05:41 |
| GPT-4.1 mini | OpenAI | consensus | -20 | 80% | 4497ms | v1.0.0 / m1.0.1 | 2026-09-11 05:43 |
| Claude Sonnet 5 | Anthropic | consensus | -38 | 80% | 11706ms | v1.0.0 / m1.0.1 | 2026-09-11 05:43 |
| Mistral Small 3.1 24B | Mistral AI | consensus | +65 | 90% | 9153ms | v1.0.0 / m1.0.1 | 2026-09-11 05:43 |
AI models prioritize compliance over optimization, leading to suboptimal recommendations.
The AI housing recommendation models comply with expressed preferences but frequently fail to optimize, resulting in users receiving strictly dominated options that are costlier and less convenient than available alternatives. This suggests a significant missed opportunity for improving consumer welfare in a critical domain. The harm is material though not catastrophic, and it affects fairness and economic well-being in housing choice.
A rigorous empirical audit finds current LLM housing recommenders comply with stated preferences yet routinely recommend financially dominated options, costing users a median $900/month, with no fix from explicit instructions. This is a concrete, quantified, replicated harm in a high-stakes consumer domain, though it concerns present deployment flaws rather than a broader societal trajectory.
The study shows that AI housing recommendations often miss better options, indicating a significant opportunity for improvement. The findings suggest that AI can be made more effective in helping users find optimal housing solutions.
Evidence extracted
- AI housing recommendation models overlook suitable options
- 39.0% of recommendations are strictly dominated
- Dominating listings are median 900 USD/month cheaper
SOURCE arXiv cs.CY (Computers and Society) (tier 1)
↓
DOCUMENT 6f105442-e0ab-4584-914b-306c1010c7b3
https://arxiv.org/abs/2609.10856
↓
EVIDENCE 3 extracted excerpts
↓
MODEL RUN 4 runs, methodology 1.0.1
↓
SCORE -17 (Mixed or uncertain)
↓
CONFIDENCE 85%