In short
- Every result links to your article and carries your outlet’s name or domain. We never return the full text.
- Customers see your headline, your own description and, at most, one excerpt of up to 25 words (two in deep mode).
- We honor robots.txt and AI opt-out signals. A block takes effect within an hour, and we delete what we stored within 72 hours.
- We don’t train AI models, and we don’t sell your content or datasets built from it.
- Want out, or want more? Use the request form. We never ask why.
Who we are, and how we send you readers
typesearch is a search API. Developers use it to give their AI agents, such as research assistants, newsroom tools and monitoring apps, a way to find recent news. When an agent searches, it gets a list of articles, each with its headline, your outlet’s name, the date and a link to your page.
Every result links to your article, and anyone who wants the full story has to open it on your site. We are a search engine, not a republisher: we don’t rewrite your articles, and we never pass on their full text.
What we take and what we show
To build the index, our crawler reads your feeds, news sitemaps, homepage and section pages. To answer a customer’s search, it reads the few articles that might match, to check that they’re relevant.
For each article, our customers get:
- your headline (up to 240 characters), the description you publish in your feed or page metadata (up to 300 characters), the date, your outlet’s name and a link to your article;
- at most one excerpt of up to 25 words (two in deep mode), never taken from the first paragraph. Articles with less than 1,200 characters of text get no excerpt at all;
- and each customer gets at most three different excerpts of the same article per day, to keep repeated searches from adding up to your article.
They never get the full text of your article, or a summary written by AI. A language model reads the text only to score how relevant it is to the search and to label it, for example by tone; it never writes about it.
Our customers are bound too. Under our Terms of Service, when they show results to people they must show your outlet’s name and link to your article, and they may not rebuild or republish your articles from excerpts, keep excerpts or descriptions for more than 30 days, or use results to train AI models.
What we store, and for how long
- The index keeps the link, headline, description, date and section of each article, for up to 12 months.
- Article text is kept for 2 days at most, only so we don’t fetch the same article twice. It never goes into our backups.
- Relevance judgments, with up to two excerpts per article, are kept for 8 days, so the same search doesn’t need the model again.
- The language model receives article text only to judge it against a customer’s search. We don’t train models with it, and we don’t sell it or datasets built from it.
What we never do
- Return the full text of your articles.
- Write AI summaries of your articles.
- Train AI models with your content, or sell it or datasets built from it.
- Use the text of paid or login-only pages.
- Get around a block. After a 401 or 403 we stop, and we never retry with another tool: no browser, no other address. A bot check (“Just a moment…”) stops us right there. After a 429 or 503 we wait, honoring
Retry-After, and retry later with the same identity. - Hide who we are. Every request names our crawler and links to its page.
Your controls
You don’t need to write to us to control what we do. We honor all of these:
- robots.txt. We have two user agents:
TypesearchBotbuilds the index from feeds, sitemaps, homepages and section pages;Typesearch-Userreads articles and uses your site’s search box when a customer’s search needs it. AUser-agent: TypesearchBotgroup applies to both; aUser-agent: Typesearch-Usergroup applies only to article reads and site search. We also honorCrawl-delay. The details are on our crawler page. - Blocks on AI search bots. If your robots.txt blocks OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot, Claude-User or Google-Extended, we treat it as a block for us too, for the same paths, unless you name TypesearchBot (or Typesearch-User) in its own group. Blocks aimed only at AI training crawlers, such as GPTBot, ClaudeBot or CCBot, don’t apply to us, because we don’t train AI models.
- Cloudflare Content Signals in robots.txt: with
search=noorai-input=no, we don’t crawl or use your site.ai-train=nohas no effect on us: we don’t train. - IETF AI preferences (
Content-Usage): withsearch=norai-use=nin robots.txt, we don’t crawl or use your site; as an HTTP header on a page (Content-UsageorContent-Signal), we don’t use that page. - TDM reservation (TDMRep): with
/.well-known/tdmrep.json, thetdm-reservation: 1header or the meta tag, we don’t read or use the content. - Robots meta tags and
X-Robots-Tag, fortypesearchbot,typesearch-useror all robots:noindexkeeps a page out of our index and results;nosnippetmeans headline and link only;max-snippet:[n]caps excerpts at n characters; text markeddata-nosnippetis never excerpted;noarchivemeans we don’t keep the text after the request;noaiornoimageaimeans we don’t use the page. - Paid or login-only content: we don’t take the text of pages marked as paid (
isAccessibleForFree: false) or behind a login. We show only the headline and the description you publish in your feed. - A reservation notice such as “Prohibida su reproducción”, on sites from Mexico or Colombia: no excerpts, only the headline, your own description and the link.
What happens when you opt out
- Within an hour of a change to your robots.txt, Content Signals or tdmrep.json, we stop fetching what they close, and articles from there stop appearing in results. Tags on a page (meta robots,
X-Robots-Tag,Content-Usage) take effect the next time we read that page; to keep pages out before that, use robots.txt or the form. - Within 72 hours we delete the data we stored from your site, usually within minutes. Backups rotate within 7 days.
- We keep a record of your choice (your domain, the date and the reason) so that we keep honoring it and, as proof that we did, a log of our requests to your site and the versions of your robots.txt, never the content of your articles, for up to two years. The contact details of whoever made the request are handled as our Privacy Policy says.
- Our Terms also require our customers to stop using stored results from a publisher that opted out within 30 days of our notice.
Tell us what you need
Use this form for anything about your site: to leave, to show headlines and links only, to remove pages, to slow our crawler down or to talk about more. We never ask why, and we don’t argue.
- You get an acknowledgment by email right away, with a case number.
- Exclusions are applied within minutes; we commit to under 24 hours.
- Stored data is deleted within 72 hours.
- You get written confirmation when it’s done.
- If the email you use is at your domain, you can confirm your request with one click. If it isn’t, we apply it anyway; for more than three sites, we check the rest first, within 24 hours.
Mentioned in the news? Choose “Remove results about me”. We review those requests case by case, as privacy law requires: see People in the news in our Privacy Policy.
Prefer email? Write to publishers@typesearch.ai.
How to verify our crawler
Every request says who it is and links to our crawler page. Without a browser, the user agents are:
Mozilla/5.0 (compatible; TypesearchBot/1.0; +https://typesearch.ai/bot)
Mozilla/5.0 (compatible; Typesearch-User/1.0; +https://typesearch.ai/bot)Our crawler’s requests come from the addresses published at https://typesearch.ai/bot/ips.json: today, a single IPv4 address, 178.128.144.188. Each one is also signed with Web Bot Auth, and our public keys are at https://api.typesearch.ai/.well-known/http-message-signatures-directory. Reverse DNS verification isn’t available yet. If something uses our name from another address, or without our signature, it isn’t us: tell us at bot@typesearch.ai.
Licensing and partnerships
If you want more — your premium content under a license, longer excerpts, a partnership — let’s talk. Choose “Licensing or partnership” in the form, or write to publishers@typesearch.ai.
Contacts
- Publishers: publishers@typesearch.ai
- Journalists: press@typesearch.ai
- Legal and copyright notices: copyright@typesearch.ai. See how to send a copyright notice.
- People mentioned in the news: privacy@typesearch.ai
- Technical questions about our crawler: bot@typesearch.ai