Applicants are asked to declare whether a model wrote their work. So this is the same question asked of the institutions: have they written down whether models may read theirs? It is a public file at a fixed address, so it can just be looked up.
22 of 370
universities whose robots.txt could be read name an AI crawler and
block it. 348 do not mention one anywhere. Read on 2026-08-24 from the QS top 500.
| Institution | Country | Crawlers blocked |
|---|---|---|
| École Polytechnique Fédérale de Lausanne | Switzerland | 8 |
| Seoul National University | Republic of Korea | 5 |
| Institut Polytechnique de Paris | France | 5 |
| New York University (NYU) | United States of America | 10 |
| The University of Nottingham | United Kingdom | 17 |
| Technische Universität Berlin | Germany | 19 |
| University of Cape Town | South Africa | 2 |
| University of Maryland, College Park | United States of America | 3 |
| University of Lausanne | Switzerland | 16 |
| Chulalongkorn University | Thailand | 1 |
| University of Porto | Portugal | 8 |
| University of Waikato | New Zealand | 5 |
| University of Johannesburg | South Africa | 13 |
| Goethe Universität Frankfurt am Main | Germany | 1 |
| University of Leicester | United Kingdom | 8 |
| Lincoln University | New Zealand | 8 |
| Dublin City University | Ireland | 1 |
| Université de Strasbourg | France | 1 |
| National Research University Higher School of Economics (HSE, Moscow) | Russian Federation | 1 |
| Universität Konstanz | Germany | 18 |
| University of California, Santa Cruz | United States of America | 9 |
| University of Jyväskylä | Finland | 1 |
The QS top 500 was joined to a public list of university domains, and a row was dropped rather
than guessed when no domain matched, which is why the set is 460 and not 500. Each
site was asked for its home page and its robots.txt once, on 2026-08-24. An institution
counts as blocking an operator only when a group naming that operator's user-agent disallows the
whole site. A partial rule is not a block, and a wildcard User-agent: * rule is not
counted at all, because it is not a statement about AI.
Two limits worth stating plainly. Silence is not consent: most of these files predate the crawlers and nobody has gone back to them. And a site can refuse crawlers at its network edge without writing anything in robots.txt, which this cannot see. So the honest finding is narrow: almost no university has written the rule down either way.
The probe is a single file and the raw output is published beside it, so anyone can re-run it and get a different answer tomorrow. That is the point of dating it.