Sitemap: https://www.sololearn.com/sitemapindex.xml # Default position for every crawler not named below: search yes, AI training no. # (This was written by Cloudflare's managed robots.txt; the file is now owned here so # that the scoped training grant further down can coexist with it. Per-agent groups # below override this for the agents they name.) User-agent: * Content-Signal: search=yes,ai-train=no,use=reference Disallow: /onboarding Disallow: /Onboarding Disallow: /*/onboarding Disallow: /*/Onboarding Disallow: /profile Disallow: /Profile Disallow: /*/profile Disallow: /*/Profile Disallow: /subscriptions Disallow: /Subscriptions Disallow: /*/subscriptions Disallow: /*/Subscriptions Disallow: /certificate Disallow: /Certificate Disallow: /*/certificate Disallow: /*/Certificate Disallow: /certificates Disallow: /Certificates Disallow: /*/certificates Disallow: /*/Certificates Disallow: /notified-* Disallow: /Notified-* Disallow: /*/notified-* Disallow: /*/Notified-* Disallow: /compiler-playground/* Disallow: /*/compiler-playground/* Disallow: /ru/compiler-playground/robots.txt Disallow: /*/compiler-playground/robots.txt # --- Generative-AI training crawlers ----------------------------------------- # Sololearn's own authored pages (home, catalog, the 86 course landings, the guide # articles, /en/plans, llms.txt) may be used for model training. Those are the facts # we want models to hold: the alternative is not "models know nothing about # Sololearn", it is "models learn Sololearn from stale third-party articles", # which is where the wrong pricing and the "it's basically a free trial" # misconception come from. Retrieval bots (OAI-SearchBot, Claude-SearchBot, # PerplexityBot, Bingbot, ChatGPT-User) were never blocked and are unaffected. # # Discuss is excluded. It is user-generated: consent to post on Sololearn is not # consent to be included in a training corpus. This also keeps ~89% of the # crawlable URL surface out of these crawls. # # CCBot and Bytespider are deliberately NOT in this group and stay blocked at the # CDN: Common Crawl redistributes far beyond the operator being granted access, # and Bytespider ignores crawl limits. # # Per-user-agent groups do NOT inherit from `User-agent: *`, so every sensitive # path is repeated here. Content-Signal is per-group: `*` keeps ai-train=no as the # default reservation of rights, and this group is the explicit, scoped grant. User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: meta-externalagent Content-Signal: search=yes,ai-train=yes,use=reference Disallow: /*/Discuss Disallow: /*/discuss Disallow: /Discuss Disallow: /discuss Disallow: /onboarding Disallow: /Onboarding Disallow: /*/onboarding Disallow: /*/Onboarding Disallow: /profile Disallow: /Profile Disallow: /*/profile Disallow: /*/Profile Disallow: /subscriptions Disallow: /Subscriptions Disallow: /*/subscriptions Disallow: /*/Subscriptions Disallow: /certificate Disallow: /Certificate Disallow: /*/certificate Disallow: /*/Certificate Disallow: /certificates Disallow: /Certificates Disallow: /*/certificates Disallow: /*/Certificates Disallow: /notified-* Disallow: /Notified-* Disallow: /*/notified-* Disallow: /*/Notified-* Disallow: /compiler-playground/* Disallow: /*/compiler-playground/* Disallow: /users Disallow: /*/users Allow: / # --- Blocked outright ------------------------------------------------------------ # CCBot: Common Crawl redistributes far beyond any operator we choose to grant. # Bytespider: ignores crawl limits. Amazonbot / CloudflareBrowserRenderingCrawler: # kept blocked for parity with the previous CDN-managed policy; revisit deliberately. User-agent: CCBot User-agent: Bytespider User-agent: Amazonbot User-agent: CloudflareBrowserRenderingCrawler Disallow: /