Zephr / Legal
Docs crawler
Synapsic LLC operates zephr-docs-crawler. It builds a citation index of public documentation. It does not train models on that corpus. The public catalog stays dark until counsel signs each source.
User-Agent
Declared token: zephr-docs-crawler. Full header includes this page as contact URL.
zephr-docs-crawler/1.0 (+https://zephr.ai/legal/crawler)
Contact:crawler@zephr.ai.
What the crawler does
Public docs in the curated catalog, preferring llms.txt. robots.txt is honored fail-closed.
- Fetch allowlisted documentation after a recorded legal posture.
- Honor robots.txt for zephr-docs-crawler (fallback *). Disallow means no crawl.
- Store a SHA-256 of the robots body on the catalog row.
- Index for retrieval and citation (snippet + URL + license).
What it does not do
No UA rotation, no paywall bypass, no training, no full public republish.
- Bypass robots, login walls, or noindex.
- Crawl red-posture or unsigned sources.
- Use crawled text to train or fine-tune models.
- Dump full third-party documentation as a public catalog.
The public catalog is dark
Unsigned greens are amber. Active crawl allowlist is empty until counsel sets reviewedBy.
Four checks before serve
- robots.txt fetched and honored.
- ToS reviewed with a recorded verdict.
- License recorded (docs license, not just repo SPDX).
- Source-level takedown path live.
First three candidates (unsigned): React, TypeScript, Playwright. Stripe, Prisma, and Supabase stay red.
How to opt out
Disallow the user-agent in robots.txt, or email a takedown.
User-agent: zephr-docs-crawler
Disallow: /
Already indexed? Copyright and DMCA or crawler@zephr.ai.
Last reviewed 2026-08-27 · operator Synapsic LLC