Tax SlopWhich universities tell AI crawlers to stay out

Applicants are asked to declare whether a model wrote their work. So this is the same question asked of the institutions: have they written down whether models may read theirs? It is a public file at a fixed address, so it can just be looked up.

22 of 370

universities whose robots.txt could be read name an AI crawler and block it. 348 do not mention one anywhere. Read on 2026-08-24 from the QS top 500.

The ones that block

InstitutionCountryCrawlers blocked
École Polytechnique Fédérale de LausanneSwitzerland8
Seoul National UniversityRepublic of Korea5
Institut Polytechnique de ParisFrance5
New York University (NYU)United States of America10
The University of NottinghamUnited Kingdom17
Technische Universität BerlinGermany19
University of Cape TownSouth Africa2
University of Maryland, College ParkUnited States of America3
University of LausanneSwitzerland16
Chulalongkorn UniversityThailand1
University of PortoPortugal8
University of WaikatoNew Zealand5
University of JohannesburgSouth Africa13
Goethe Universität Frankfurt am MainGermany1
University of LeicesterUnited Kingdom8
Lincoln UniversityNew Zealand8
Dublin City UniversityIreland1
Université de StrasbourgFrance1
National Research University Higher School of Economics (HSE, Moscow)Russian Federation1
Universität KonstanzGermany18
University of California, Santa CruzUnited States of America9
University of JyväskyläFinland1

How this was measured

The QS top 500 was joined to a public list of university domains, and a row was dropped rather than guessed when no domain matched, which is why the set is 460 and not 500. Each site was asked for its home page and its robots.txt once, on 2026-08-24. An institution counts as blocking an operator only when a group naming that operator's user-agent disallows the whole site. A partial rule is not a block, and a wildcard User-agent: * rule is not counted at all, because it is not a statement about AI.

Two limits worth stating plainly. Silence is not consent: most of these files predate the crawlers and nobody has gone back to them. And a site can refuse crawlers at its network edge without writing anything in robots.txt, which this cannot see. So the honest finding is narrow: almost no university has written the rule down either way.

The probe is a single file and the raw output is published beside it, so anyone can re-run it and get a different answer tomorrow. That is the point of dating it.

All 460