what is robots.txt? the file that tells crawlers where they may go, in plain words
seo · Mar 19, 2026 · 5 min read
robots.txt is a plain text file that lives at the root of your domain — mybusiness.com/robots.txt — and tells search engine crawlers which parts of the site they are allowed to crawl. it is the bouncer's guest list: googlebot arrives, reads the file first, and follows the rules it finds.
what the file actually says
the syntax is two lines repeated, and nothing more:
User-agent: Googlebot
Disallow: /admin/
"User-agent" names the crawler ("*" means all of them), "Disallow" lists path prefixes it should skip. that is the whole vocabulary. most sites either have an empty ruleset — crawl everything — or a single line blocking an admin path or a search-results page.
the file exists for one job: saving crawlers from wasting visits on pages that have no value in search. a large site uses it to keep bots out of faceted filters and internal logs; a small site almost never needs the complexity.
what robots.txt cannot do
this is where the misconceptions live, so the hard limits, plainly:
-
it does not hide a page from google. a disallowed page can still appear in results — without its content — if other pages link to it. the only way to keep a page out of the index is a
noindexdirective on the page itself, which requires the crawler to be allowed to visit it. blocking and noindexing are opposites, and confusing them is the classic mistake. -
it is not security. a disallowed path is a published announcement that the path exists. anything that must stay private needs authentication, not a note asking bots to look away.
-
it is not a ranking tool. disallowing pages does not concentrate "power" anywhere; it only removes crawl visits.
what you should actually do
-
look at yours once. open
yourdomain/robots.txt. it exists on almost every platform by default. if it says nothing scary, nothing to do. -
make sure it does not block css or js. an ancient copy-paste rule blocking
/assets/makes google render your pages half-blind, which quietly hurts how the pages are evaluated. this is the only common way robots.txt actively damages a small site. -
point it at your sitemap with one
Sitemap:line — the handshake described in do i need an xml sitemap. -
watch its effects in google search console, which reports crawl anomalies by name.
FAQ: the questions i actually get
do i need to create a robots.txt?
probably not. builders and hosts ship a sane default. create one the day you have a specific reason: a crawl-heavy parameter maze, a staging area, an admin path generating noise in crawl stats.
can i block my staging site with robots.txt?
you can block crawling, but the url can still leak and be indexed as an empty shell. staging sites belong behind a password. the password is the fix; robots.txt is a courtesy notice.
does robots.txt affect my ranking?
only indirectly, and only in the negative: blocking things google needs (css, js, images) degrades how pages render and evaluate. used sanely, it is invisible to rankings — the goal of how do i get my website on google is being crawled well, not less.