Standards
What AI research gets wrong about Iran.
“Couldn’t an AI assistant produce this?” is a fair question, and it deserves a measured answer rather than a marketing one. So we measured it: general-purpose AI research, with live web access, against this desk’s registry-verified record — errors documented, method published, our own misses included.
9 / 100Provable errors across 100 field-answers — 4 of them critical
8 / 20Companies whose registered operating entity the baseline could not name
3Errors the same run found in our own record — corrected and logged the same day
01Method
July 2026, run 01.Twenty companies from our public coverage — flagships and verification-hard cases — across five fields each: current chief executive, founding year, ownership, registered operating entity, and operational status. The baseline: independent AI research agents with live web search, instructed to research as a diligent assistant would, to cite a source for every claim, to say “not found” rather than guess — and barred from citing Tehran Index. Their hundred answers were then scored against our record, with one deliberate asymmetry: an error only counts against the baseline where a registry-grade source proves the matter. Where the baseline merely disagreed with us, the item went to desk review instead — because it might be us who was wrong.
02Critical findings
Four answers were wrong in ways that would mislead a diligence process — each delivered fluently, with citations attached.
| Company | Field | The AI baseline | The verified record |
|---|---|---|---|
| Fidibo | Legal entity | Named the operating company with a registry identifier that returns zero results in the corporate registry — a wrong ID circulating on aggregator sites, repeated as registry fact. | The registered entity behind the brand is a different company, confirmed against the registry. (Identifiers are withheld here by policy.) |
| Jobinja | Status | "Active" — stated at high confidence because the website is live and serving listings. | The operating company is registered as dissolved, in liquidation. A live website is not a live company. |
| SnappPay | Managing director | Named the chairman as "founder and CEO", sourced from a LinkedIn self-description. | The registered managing director is a different person, per the corporate registry — appointed February 2023. The chairman and the registered MD are different roles held by different people. |
| Digikala | Ownership | Reproduced the pre-2024 shareholder structure from an encyclopedia page. | Misses the ~40% acquisition by MCI/Hamrah-e Aval in 2024 — the largest ownership event in Iranian e-commerce. |
03Material findings
Cafe Bazaar — the completed January 2025 change of control was invisible to the baseline ("post-2025 ownership not found").
Digikala — leadership answer was the former co-founder CEO, not the current group CEO.
Jobvision — found one of the two co-CEOs; the second, board-confirmed against the registry, was "not found".
YektaNet — attached the wrong registered entity from a registry aggregator.
04The identity gap
The most consistent failure was structural, not factual: for eight of the twenty companies, the baseline could not name the registered legal entity that actually operates the brand — including some of the most prominent consumer platforms in the country. That layer — which entity you would actually contract with, diligence, or trace — is close to invisible to web-scale research, because it lives in Persian-language registry records behind access barriers. It is precisely the layer this desk anchors every profile to.
05Why it fails
Watching the baseline work explains the errors. Iranian primary sources — the corporate registry, the major Persian tech press — were repeatedly unreachable to it, so it fell back on English-language aggregators and mirror sites. Aggregators echo each other, including each other’s mistakes; the wrong Fidibo identifier exists on exactly such a site, which is how it arrived in the baseline’s answer wearing a citation. The failure mode that matters is not ignorance — it is confident, cited, wrong.
06What it gets right
Credit where due: the baseline handled well-documented, English-visible facts cleanly — Tapsi’s listing and change of control, Myket’s operator and group, Filimo’s structure, founding years across most of the sample. General AI research is genuinely good at what the open web already says clearly. The errors concentrate exactly where the open web goes quiet: registry identity, current officers, recent transactions, and whether a company behind a live website still legally exists.
07The same test, applied to us
A benchmark that cannot cost us anything would not be worth publishing. The run’s disagreements sent five items to desk review, and on three of them the outside view was ahead of us: two chief-executive fields were stale and one independence label was wrong. All three were re-verified against primary sources and corrected the same day, and the corrections are logged in the open on the corrections page. That loop — measure, correct, publish, repeat — is the standard this desk holds itself to, and the benchmark re-runs quarterly under the same method.
08Limits
This is one run, at one date, with one configuration; results will vary by tool and by day, and the tools improve. We benchmark the category of general-purpose AI research, not any vendor, and we expect the gap to move — which is why the measurement repeats. What a re-run cannot change is the structural finding: research on this market is only as good as its access to primary sources, and its willingness to reject what fails verification.