The State of the Evidence
How much of the smart-city record has actually been examined by anyone outside the organisation that built it — counted across all 651 source-checked entries in this atlas.
Every entry in this atlas carries a grade for the strongest source behind it. This page counts those grades instead of the projects — it is the only page here that is not about what cities built, but about how much anyone actually knows about what they built.
The reason to count this is simple. Almost everything written about smart cities is written by someone with a stake in the answer: the vendor that sold the system, the department that approved it, the consultancy that advised on both. Those accounts are not worthless, but they are not measurements either, and nothing in the usual literature tells you the ratio. Grading 651 projects one at a time makes the ratio visible.
✔ Every published entry is checked against the public sources it cites before publication.
What is being counted
Each published entry gets one evidence grade: the strongest kind of source that supports its core claims — not a grade per figure, and not a judgement of whether the project was any good. A well-run project whose results only its operator has published is official. A failure documented by a national audit office is independent. The grade describes the state of the record, not the state of the city.
The grades are assigned under a written protocol during the source check, and they are AI-assigned like the rest of the atlas — the entries are AI-researched summaries of the sources each page lists, produced under a documented rule set rather than individually re-checked by a human. That is worth stating plainly on a page about evidence quality: this is a consistent reading of the record by one method, not an audit by a third party. The underlying data is open, so the reading can be checked.
- 🔬 Peer-reviewed
- Impact confirmed by an independent peer-reviewed or academic evaluation.
- 📊 Independent
- Impact confirmed by an independent non-academic evaluation (audit, EU/government review, NGO).
- 🏛️ Official
- Impact figures self-reported by the operator or city (official report or press release).
- 📰 Media
- Impact reported by the press only; no primary evaluation located.
- • No measured figure
- No independently verifiable impact figure located; the entry rests on qualitative description.
1 🏛️ Two-thirds of the record is the operator's own word
Across all 651 published entries, 68.0% rest on figures published by the body that runs the project — the city, the utility, the transport authority, the vendor. Peer-reviewed evaluation exists for 8.4%. Add every independent audit, regulator's review and NGO evaluation to the academic work and the share of the atlas that anyone outside the project has examined is 20.1%.
| Strongest source behind the entry | Entries | Share |
|---|---|---|
| 🔬 Peer-reviewed | 55 | 8.4% |
| 📊 Independent | 76 | 11.7% |
| 🏛️ Official | 443 | 68.0% |
| 📰 Media | 52 | 8.0% |
| • No measured figure | 25 | 3.8% |
This is not an accusation of dishonesty. Operators publish the numbers they have, and often they are the only body positioned to collect them. It does mean that when a city compares itself against what other cities report, it is mostly comparing press releases.
2 📊 Where a result is claimed, it is usually the seller's number
375 entries state a concrete outcome figure — a percentage saved, a journey time cut, a leak rate reduced. Those are the numbers that travel: into slide decks, procurement documents and news copy. 67.5% of them come from the operator. 11.2% have been through peer review.
| Strongest source behind the figure | Entries | Share of the 375 with a figure |
|---|---|---|
| 🔬 Peer-reviewed | 42 | 11.2% |
| 📊 Independent | 44 | 11.7% |
| 🏛️ Official | 253 | 67.5% |
| 📰 Media | 36 | 9.6% |
The atlas labels this on every project page rather than smoothing it away, which is why the count is possible at all. A figure with an operator's name on it is reported as an operator's figure, never as a measured fact.
3 🔍 The recorded failure rate depends on who was allowed to measure
Sort the atlas by who examined a project and the failure rate moves sharply. Among entries whose best source is the operator, 5.4% are recorded as failed or scaled back. Among entries examined by an independent auditor, regulator or NGO, 21.1% — roughly four times as many. Where the press is the only source that looked, 23.1%.
| Strongest source behind the entry | Entries | Failed or scaled back | Rate |
|---|---|---|---|
| 🔬 Peer-reviewed | 55 | 5 | 9.1% |
| 📊 Independent | 76 | 16 | 21.1% |
| 🏛️ Official | 443 | 24 | 5.4% |
| 📰 Media | 52 | 12 | 23.1% |
| • No measured figure | 25 | 1 | 4.0% |
Read the peer-reviewed row before drawing the obvious conclusion: at 9.1% it sits below the operator-reported rate, which is the opposite of the pattern. That is a selection effect, and this corpus can show it rather than just claim it — the peer-reviewed entries have been running for a median of 11 years against 8 for the operator-reported ones, and 90.9% of them are still live. Academics study systems that lasted long enough to be worth a multi-year study and a journal cycle; the pilot that folded after eighteen months never becomes a paper. The academic subset is not a random sample of the atlas, it is a sample of survivors, and at 55 entries a handful of cases moves it. So the claim the table supports is narrower than "scrutiny finds failure": among the groups that produced a figure at all, the operator-reported share of failure is the lowest, and independent and press examination both land several times higher. Whether self-reporting omits shortfalls, or outside evaluations simply get commissioned once a project is contested, this data cannot say. Both are ordinary. Both are reasons to grade the source.
4 🕒 The newest projects are the least examined
Evaluation lags deployment, and the gap is wide. Of the entries that started before 2010, 26.4% have been examined by someone at arm's length. For projects launched since 2020, 13.7%.
| Project started | Entries | Examined at arm's length | Share |
|---|---|---|---|
| Before 2010 | 106 | 28 | 26.4% |
| 2010–2014 | 136 | 31 | 22.8% |
| 2015–2019 | 213 | 46 | 21.6% |
| 2020 onward | 190 | 26 | 13.7% |
This is the least surprising finding here and possibly the most consequential. A city choosing technology in 2026 is choosing from the cohort with the thinnest evidence base, and the systems with a decade of independent study behind them are the ones now being replaced. The evidence arrives, on average, after the procurement decision.
5 💸 Four out of five entries have no public price
A cost figure could be traced for 19.2% of published entries — 125 of 651. The rest are public infrastructure whose price is not in any source we could find. Where money is visible, it is visible unevenly: 40.4% of public-private partnerships carry a cost figure, against 6.9% of privately run systems.
| Funding model | Entries | With a cost figure | Share |
|---|---|---|---|
| Public-private partnership | 47 | 19 | 40.4% |
| Utility-funded | 49 | 11 | 22.4% |
| Mixed / blended | 41 | 9 | 22.0% |
| Public budget | 375 | 75 | 20.0% |
| Grant-funded | 37 | 5 | 13.5% |
| Private | 29 | 2 | 6.9% |
The pattern is not that PPPs are more open by disposition; it is that a contract award has to be published and an internal budget line does not. Publication follows procurement law, not policy on transparency.
6 🗂️ The evidence is thinnest where the spending is heaviest
Independent examination is not spread evenly across the field. Mobility is the largest category in the atlas at 252 entries and one of the least independently examined; energy is the thinnest of all, with 1.7% peer-reviewed.
| Category | Entries | Peer-reviewed | Examined at arm's length |
|---|---|---|---|
| 👥 Society | 9 | 3 | 4 (44.4%) |
| 🌱 Environment | 96 | 15 | 29 (30.2%) |
| 📡 Data & IoT | 130 | 8 | 27 (20.8%) |
| 🏛️ Governance | 102 | 8 | 21 (20.6%) |
| 🚲 Mobility | 252 | 19 | 42 (16.7%) |
| ⚡ Energy | 60 | 1 | 7 (11.7%) |
Environmental projects score best, which is less a compliment to city departments than a side effect of academia: air quality, heat and flooding are established research subjects with people already measuring them for other reasons. Where no research community happens to be watching, nobody is.
What this report is not
- Not a sample of the world's smart-city projects. This atlas is curated. Entries are selected for being documentable, which biases the corpus toward projects that left a public trail — and away from the ones that quietly never happened. Every share here describes this record, not the field.
- Not a quality ranking. The evidence grade says who examined a project, not how well it worked. Some of the best entries in the atlas are graded official because nobody outside the operator has ever looked.
- Failure counts are a floor. An entry is only recorded as failed or scaled back when a named source says so. Quiet abandonment rarely produces one, so the true share is higher than 8.9% — by how much, we cannot say.
- "Ongoing" is a default, not a verdict. 69.7% of entries are recorded as ongoing, which means running with no clear outcome yet. It should not be read as success and does not belong in a success rate.
- The grader is not independent of the graded. The same AI-assisted process that wrote the entries assigned the grades. That is a real limitation, stated here rather than in a footnote — the defence is that the rule set is written down, the sources are on every page, and the data is open enough for anyone to grade it differently.
Check this yourself
Nothing on this page is a private calculation. The whole corpus is published as JSON under CC BY 4.0, and every figure above is a count over two fields — evidence and outcome — plus cost, cat, fundingModel and yearStart. Download the data, group by the field, and you will get these numbers or find a mistake worth telling us about.
Common questions
What does "evidence: official" actually mean?
That the strongest source supporting the entry is the operator or the city itself — an annual report, an evaluation the department commissioned, a press release with figures. It is the most common state of the record and it is not the same as an unsupported claim: the source exists, is named on the project page, and is public. It has simply not been checked by anyone outside the project.
Does a low share of peer-reviewed evidence mean smart cities do not work?
No. It means most of them have not been independently measured, which is a different and less satisfying statement. Some are certainly working; the point is that in most cases nobody outside the organisation running the system has verified it, so a city comparing options is mostly comparing self-reported results.
Why is the failure rate higher for independently evaluated projects?
We cannot tell from the data whether independent scrutiny finds shortfalls that self-reporting omits, or whether evaluations get commissioned more often when a project is already contested. Both are plausible and both are common. The finding is the association, not a cause.
Who graded the evidence, and can that be trusted?
The grades were assigned during the source check, by the same AI-assisted editorial process that produces the entries, following a written rule set. That process is not independent of the material it grades, which is why the underlying data is published openly under CC BY 4.0 and every entry lists the sources it was graded on. Disagreeing with a specific grade is easy and welcome — the contribute form takes corrections.
How often does this report change?
It is recomputed from the corpus on every publication, so it always describes the atlas as it stands today rather than a fixed snapshot. As more entries are added and existing ones re-verified, the shares will move — including, we would hope, upward.
Where to read next
The findings above have addresses elsewhere in the atlas: every entry recorded as failed or scaled back is listed on smart city project failures, grouped by how it went wrong; why smart cities fail argues the pattern behind them; and how cities are ranked takes apart the indices that turn self-reported figures into league tables.
Recomputed on every build from the published corpus, so the figures on this page are never older than the atlas itself. Last publication: .
