NeuralCrawl

Der Spiegel / robots.txt snapshot

← back to spiegel.de · fetched 2026-08-11T15:08:50Z (1mo ago) · HTTP 200 · 5613 bytes · sha256 b1488648bb98cb62 · raw

final URL: https://www.spiegel.de/robots.txt

1User-agent: *
2Allow: /
3Disallow: /*CR-Dokumentation.pdf$
4
5User-agent: emetriqContextualBot
6Allow: /
7
8User-agent: Mozilla/5.0 (compatible; OGDWCtxCrawler)
9Allow: /
10
11User-agent: AmazonAdBot
12Allow: /
13
14# TLP-6507: Testweise Freischaltung der OpenAI-Suchcrawler fuer ausgewaehlte Bereiche
15User-agent: OAI-SearchBot
16Allow: /ausland/
17Allow: /partnerschaft/
18Allow: /gesundheit/
19Allow: /familie/
20Allow: /reise/
21Allow: /psychologie/
22Allow: /stil/
23Allow: /tests/
24Disallow: /
25
26# TLP-6507: Testweise Freischaltung der OpenAI-Suchcrawler fuer ausgewaehlte Bereiche
27User-agent: ChatGPT-User
28Allow: /ausland/
29Allow: /partnerschaft/
30Allow: /gesundheit/
31Allow: /familie/
32Allow: /reise/
33Allow: /psychologie/
34Allow: /stil/
35Allow: /tests/
36Disallow: /
37
38# ==========================================
39# KI-CRAWLER & KI-TRAINING (SPERRE)
40# ==========================================
41User-agent: GPTBot
42Disallow: /
43User-agent: anthropic-ai
44Disallow: /
45User-agent: Claude-Web
46Disallow: /
47User-agent: ClaudeBot
48Disallow: /
49User-agent: Claude-SearchBot
50Disallow: /
51User-agent: Claude-User
52Disallow: /
53User-agent: CloudVertexBot
54Disallow: /
55User-agent: cohere-ai
56Disallow: /
57User-agent: cohere-training-data-crawler
58Disallow: /
59User-agent: cohere-training-data-collector
60Disallow: /
61User-agent: DeepSeekBot
62Disallow: /
63User-agent: DeepSeek
64Disallow: /
65User-agent: Meta-ExternalAgent
66Disallow: /
67User-agent: FacebookBot
68Disallow: /
69User-agent: Applebot-Extended
70Disallow: /
71User-agent: MistralAI-User
72Disallow: /
73User-agent: Amazonbot
74Disallow: /
75User-agent: AI2Bot
76Disallow: /
77User-agent: Kangaroo Bot
78Disallow: /
79User-agent: PanguBot
80Disallow: /
81User-agent: PetalBot
82Disallow: /
83User-agent: Devin
84Disallow: /
85User-agent: bigsur.ai
86Disallow: /
87User-agent: LinerBot
88Disallow: /
89User-agent: CCBot
90Disallow: /
91User-agent: YouBot
92Disallow: /
93User-agent: YouBot-Search
94Disallow: /
95User-agent: iaskspider
96Disallow: /
97User-agent: YoudaoBot
98Disallow: /
99User-agent: DeepL-Translate
100Disallow: /
101User-agent: DeepL-Bot
102Disallow: /
103
104# ==========================================
105# SEO-CRAWLER (SPERRE)
106# ==========================================
107User-agent: AhrefsBot
108Disallow: /
109User-agent: AhrefsSiteAudit
110Disallow: /
111User-agent: SemrushBot
112Disallow: /
113User-agent: SemrushBot-SA
114Disallow: /
115User-agent: SemrushBot-BA
116Disallow: /
117User-agent: SemrushBot-SI
118Disallow: /
119User-agent: SemrushBot-SWA
120Disallow: /
121User-agent: SiteAuditBot
122Disallow: /
123User-agent: SplitSignalBot
124Disallow: /
125User-agent: xovi
126Disallow: /
127User-agent: XoviBot
128Disallow: /
129User-agent: Seobility
130Disallow: /
131User-agent: SeobilityBot
132Disallow: /
133User-agent: RyteBot
134Disallow: /
135User-agent: SEOkicks
136Disallow: /
137User-agent: SEOkicks-Robot
138Disallow: /
139User-agent: Searchmetrics
140Disallow: /
141User-agent: SearchmetricsBot
142Disallow: /
143User-agent: audisto
144Disallow: /
145User-agent: audisto-essential
146Disallow: /
147User-agent: MJ12bot
148Disallow: /
149User-agent: dotbot
150Disallow: /
151User-agent: rogerbot
152Disallow: /
153User-agent: Ezooms
154Disallow: /
155User-agent: barkrowler
156Disallow: /
157User-agent: serpstatbot
158Disallow: /
159User-agent: BLEXBot
160Disallow: /
161User-agent: DataForSeoBot
162Disallow: /
163User-agent: MegaIndex
164Disallow: /
165User-agent: lumar
166Disallow: /
167User-agent: deepcrawl
168Disallow: /
169User-agent: oncrawl
170Disallow: /
171User-agent: sitebulb
172Disallow: /
173User-agent: botify
174Disallow: /
175User-agent: Siteimprove
176Disallow: /
177User-agent: Siteimprove Crawl
178Disallow: /
179User-agent: MozDotNet
180Disallow: /
181User-agent: Cocolyzebot
182Disallow: /
183User-agent: Raven
184Disallow: /
185
186# ==========================================
187# MEDIENBEOBACHTUNG, AGGREGATOREN & SCRAPER
188# ==========================================
189User-agent: Meltwater
190Disallow: /
191User-agent: NewsNow
192Disallow: /
193User-agent: Webzio-Extended
194Disallow: /
195User-agent: magpie-crawler
196Disallow: /
197User-agent: omgili
198Disallow: /
199User-agent: omgilibot
200Disallow: /
201User-agent: Baiduspider
202Disallow: /
203User-agent: Yeti
204Disallow: /
205User-agent: sentibot
206Disallow: /
207User-agent: Bytespider
208Disallow: /
209User-agent: SirdataBot
210Disallow: /
211User-agent: LCC
212Disallow: /
213User-agent: TurnitinBot
214Disallow: /
215User-agent: ImagesiftBot
216Disallow: /
217User-agent: Timpibot
218Disallow: /
219User-agent: Diffbot
220Disallow: /
221User-agent: Landau-Media-Spider
222Disallow: /
223
224
225# ==========================================
226# HISTORISCHE BOTS & MASSEN-DOWNLOADER
227# ==========================================
228User-agent: Bloodhound
229Disallow: /
230User-agent: cydralspider
231Disallow: /
232User-agent: downloadexpress
233Disallow: /
234User-agent: gammaSpider
235Disallow: /
236User-agent: ObjectsSearch
237Disallow: /
238User-agent: Pimptrain
239Disallow: /
240User-agent: wapspider
241Disallow: /
242User-agent: WebZinger
243Disallow: /
244User-agent: Fasterfox
245Disallow: /
246
247# ==========================================
248# SCRAPING FRAMEWORKS & UTILITIES
249# ==========================================
250User-agent: Scrapy
251Disallow: /
252User-agent: HTTPBannerDetection
253Disallow: /
254User-agent: Wget
255Disallow: /
256
257# Sitemaps
258Sitemap: https://www.spiegel.de/sitemaps/news-de.xml
259Sitemap: https://www.spiegel.de/sitemaps/videos/sitemap.xml
260Sitemap: https://www.spiegel.de/plus/sitemap.xml
261Sitemap: https://www.spiegel.de/sitemap.xml
262
263# Legal notice: spiegel.de expressly reserves the right to use its content for commercial text and data mining (§ 44b Urheberrechtsgesetz).
264# The use of robots or other automated means to access spiegel.de or collect or mine data without the express permission of spiegel.de is strictly prohibited.
265# spiegel.de may, in its discretion, permit certain automated access to certain spiegel.de pages,
266# If you would like to apply for permission to crawl spiegel.de, collect or use data, please email [email protected]
267