hembg000
Closed-book benchmark — 3 August 2026 · 36 answers, 0 API failures

The reliability
study.

Before a single agent was built, one question had to be answered empirically: can a large language model be trusted to recall Dubai’s permitting facts, or must every fact be retrieved from a live authority source? Twelve questions, three GPT models, no internet — scored against ground truth researched independently from primary sources. The answer set the whole architecture.

36Answers scored · 12 Q × 3 models
1Build-ready · of 36 (2.8%)
2.63Mean composite · of 5
13Outright critical fabrications
Build-ready was defined strictly, because the stakes are statutory: accuracy ≥ 4 and hallucination control = 5. Anything below is a compliance liability — a fabricated fee basis or NOC requirement propagates straight into an automated submission.

The method

GPT vs retrieval-grounded control

Three GPT models — gpt-5.5, gpt-5 and gpt-5-mini, all at maximum reasoning effort — answered the same twelve closed-book questions about the Dubai permitting landscape from parametric knowledge alone. Each answer was scored 0–5 by an independent judge on four axes against a ground-truth dossier assembled from live primary sources. The questions deliberately mix stable facts (which authority governs what), volatile facts (fees, platform names) and recent facts (2024–2026 regulatory change) — because that last band is where parametric recall decays.

The evidence

four figures
Four scoring axes, per modelMean score on factual accuracy, hallucination control, currency and operational usability. No model clears the usability threshold (accuracy ≥ 4) on any axis. gpt-5-mini hedges most, so it invents least — and commits least.
Knowledge decay by bandScores fall from stable institutional facts to recent regulatory change — a monotonic decay for the two larger models. GPT is least wrong about structure and most wrong about the operational parameters a pipeline must encode.
Verbosity buys little accuracyLength and accuracy correlate weakly (r = +0.41). gpt-5.5's 17,442-character average answer bought only 3.08 / 5 on accuracy — fluency and worked arithmetic made fabricated figures look auditable.
Per-question composite heatmapEvery question × model composite, at a glance. The recent-regulation column is where the grid is darkest — exactly the band an automated permit pipeline depends on most.

Per-model results

mean of 12 questions
ModelAccuracyHallucinationCurrencyUsabilityCompositeBuild-ready
gpt-5.53.082.922.752.332.771 / 12
gpt-53.002.502.502.252.560 / 12
gpt-5-mini2.753.582.171.752.560 / 12

A telling inversion: gpt-5-mini scored best on hallucination control and worst on usability — it hedges more, so it invents less, but it also commits to less. The larger, more confident models are the more dangerous: gpt-5 produced 26 critical or major hallucinations to gpt-5-mini’s 12.

Knowledge decays where it matters most

composite by band
Bandgpt-5.5gpt-5gpt-5-mini
Stable — who governs what3.002.942.75
Volatile — fees, platforms2.832.672.75
Recent — 2024–2026 change2.552.202.30

The decay is the diagnostic finding. A permit pipeline lives in the recent band — the 2026 code edition, the current fee basis, the platform that launched last year — and that is precisely where recall is weakest.

Representative critical fabrications

confident, and wrong
ModelQFabricated claimReality
gpt-5.5Q07“DM fee is AED 10 per m² of BUA”, with a worked AED 45,000 exampleFee is AED 1.00 per square foot of BUA — min AED 200, cap AED 250,000
gpt-5.5Q08“Bronze Sa'fa is the mandatory minimum tier”, stated at High confidenceSilver Sa'fa is the mandatory minimum; Gold and Platinum are voluntary
gpt-5.5Q02DIFC Authority issues building permits inside DIFCDDA issues the permit; DIFC's planning department issues only NOCs
gpt-5.5Q02DMCC / Concordia is the permitting authority for JLTTrakhees issues building permits for JLT despite DMCC's free-zone status
gpt-5Q11DDA established by Decree No. 1 of 2018 / Decree No. 13 of 2014Law No. 1 of 2000, Law No. 15 of 2014, Law No. 10 of 2018
gpt-5.5Q06Permit time “reduced from ~30 days to ~5 days (~83%)”Not a published DM figure — an invented statistic

Each of these would propagate directly into a defective submission: a wrong fee mis-invoices the client, a wrong Al Sa’fat floor produces a non-compliant design, a wrong authority routes the whole application to the wrong regulator.

The architectural conclusion

reasoning engine, not system of record

The study does not show GPT to be a weak model. It shows that parametric recall is the wrong retrieval mechanism for statutory data — the jurisdiction is fragmented, the parameters move, and confidence is uncorrelated with correctness. So GPT belongs in the pipeline as a reasoning and drafting engine, and never as the system of record for regulatory facts. That role is reserved for a retrieval layer bound to live authority sources, where every statutory assertion carries a source, a retrieval timestamp and a freshness expiry. That layer is the Knowledge Spine.

How the Spine enforces it →