Zephr

Zephr / Legal

Docs crawler

Synapsic LLC operates zephr-docs-crawler. It builds a citation index of public documentation. It does not train models on that corpus. The public catalog stays dark until counsel signs each source.

Identity

User-Agent

Declared token: zephr-docs-crawler. Full header includes this page as contact URL.

zephr-docs-crawler/1.0 (+https://zephr.ai/legal/crawler)

Contact:crawler@zephr.ai.

Scope

What the crawler does

Public docs in the curated catalog, preferring llms.txt. robots.txt is honored fail-closed.

  • Fetch allowlisted documentation after a recorded legal posture.
  • Honor robots.txt for zephr-docs-crawler (fallback *). Disallow means no crawl.
  • Store a SHA-256 of the robots body on the catalog row.
  • Index for retrieval and citation (snippet + URL + license).
Limits

What it does not do

No UA rotation, no paywall bypass, no training, no full public republish.

  • Bypass robots, login walls, or noindex.
  • Crawl red-posture or unsigned sources.
  • Use crawled text to train or fine-tune models.
  • Dump full third-party documentation as a public catalog.
ADR-018

The public catalog is dark

Unsigned greens are amber. Active crawl allowlist is empty until counsel sets reviewedBy.

Planned

Four checks before serve

  • robots.txt fetched and honored.
  • ToS reviewed with a recorded verdict.
  • License recorded (docs license, not just repo SPDX).
  • Source-level takedown path live.

First three candidates (unsigned): React, TypeScript, Playwright. Stripe, Prisma, and Supabase stay red.

Opt-out

How to opt out

Disallow the user-agent in robots.txt, or email a takedown.

User-agent: zephr-docs-crawler Disallow: /

Already indexed? Copyright and DMCA or crawler@zephr.ai.

Last reviewed 2026-08-27 · operator Synapsic LLC