NewYour Exa or Tako SDK, up to 5× cheaper.

For publishers

typesearch sends readers to your site

typesearch is a news search engine for AI agents. We index headlines and links so that developers’ agents can find your reporting, and every result points to your article. Here is exactly what we take, what we show, what we keep, and how to change any of it.

Also in: Español · Português

Last updated September 24, 2026

In short

  • Every result links to your article and carries your outlet’s name or domain. We never return the full text.
  • Customers see your headline, your own description and, at most, one excerpt of up to 25 words (two in deep mode).
  • We honor robots.txt and AI opt-out signals. A block takes effect within an hour, and we delete what we stored within 72 hours.
  • We don’t train AI models, and we don’t sell your content or datasets built from it.
  • Want out, or want more? Use the request form. We never ask why.

Who we are, and how we send you readers

typesearch is a search API. Developers use it to give their AI agents, such as research assistants, newsroom tools and monitoring apps, a way to find recent news. When an agent searches, it gets a list of articles, each with its headline, your outlet’s name, the date and a link to your page.

Every result links to your article, and anyone who wants the full story has to open it on your site. We are a search engine, not a republisher: we don’t rewrite your articles, and we never pass on their full text.

What we take and what we show

To build the index, our crawler reads your feeds, news sitemaps, homepage and section pages. To answer a customer’s search, it reads the few articles that might match, to check that they’re relevant.

For each article, our customers get:

  • your headline (up to 240 characters), the description you publish in your feed or page metadata (up to 300 characters), the date, your outlet’s name and a link to your article;
  • at most one excerpt of up to 25 words (two in deep mode), never taken from the first paragraph. Articles with less than 1,200 characters of text get no excerpt at all;
  • and each customer gets at most three different excerpts of the same article per day, to keep repeated searches from adding up to your article.

They never get the full text of your article, or a summary written by AI. A language model reads the text only to score how relevant it is to the search and to label it, for example by tone; it never writes about it.

Our customers are bound too. Under our Terms of Service, when they show results to people they must show your outlet’s name and link to your article, and they may not rebuild or republish your articles from excerpts, keep excerpts or descriptions for more than 30 days, or use results to train AI models.

What we store, and for how long

  • The index keeps the link, headline, description, date and section of each article, for up to 12 months.
  • Article text is kept for 2 days at most, only so we don’t fetch the same article twice. It never goes into our backups.
  • Relevance judgments, with up to two excerpts per article, are kept for 8 days, so the same search doesn’t need the model again.
  • The language model receives article text only to judge it against a customer’s search. We don’t train models with it, and we don’t sell it or datasets built from it.

What we never do

  • Return the full text of your articles.
  • Write AI summaries of your articles.
  • Train AI models with your content, or sell it or datasets built from it.
  • Use the text of paid or login-only pages.
  • Get around a block. After a 401 or 403 we stop, and we never retry with another tool: no browser, no other address. A bot check (“Just a moment…”) stops us right there. After a 429 or 503 we wait, honoring Retry-After, and retry later with the same identity.
  • Hide who we are. Every request names our crawler and links to its page.

Your controls

You don’t need to write to us to control what we do. We honor all of these:

  • robots.txt. We have two user agents: TypesearchBot builds the index from feeds, sitemaps, homepages and section pages; Typesearch-User reads articles and uses your site’s search box when a customer’s search needs it. A User-agent: TypesearchBot group applies to both; a User-agent: Typesearch-User group applies only to article reads and site search. We also honor Crawl-delay. The details are on our crawler page.
  • Blocks on AI search bots. If your robots.txt blocks OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot, Claude-User or Google-Extended, we treat it as a block for us too, for the same paths, unless you name TypesearchBot (or Typesearch-User) in its own group. Blocks aimed only at AI training crawlers, such as GPTBot, ClaudeBot or CCBot, don’t apply to us, because we don’t train AI models.
  • Cloudflare Content Signals in robots.txt: with search=no or ai-input=no, we don’t crawl or use your site. ai-train=no has no effect on us: we don’t train.
  • IETF AI preferences (Content-Usage): with search=n or ai-use=n in robots.txt, we don’t crawl or use your site; as an HTTP header on a page (Content-Usage or Content-Signal), we don’t use that page.
  • TDM reservation (TDMRep): with /.well-known/tdmrep.json, the tdm-reservation: 1 header or the meta tag, we don’t read or use the content.
  • Robots meta tags and X-Robots-Tag, for typesearchbot, typesearch-user or all robots: noindex keeps a page out of our index and results; nosnippet means headline and link only; max-snippet:[n] caps excerpts at n characters; text marked data-nosnippet is never excerpted; noarchive means we don’t keep the text after the request; noai or noimageai means we don’t use the page.
  • Paid or login-only content: we don’t take the text of pages marked as paid (isAccessibleForFree: false) or behind a login. We show only the headline and the description you publish in your feed.
  • A reservation notice such as “Prohibida su reproducción”, on sites from Mexico or Colombia: no excerpts, only the headline, your own description and the link.

What happens when you opt out

  • Within an hour of a change to your robots.txt, Content Signals or tdmrep.json, we stop fetching what they close, and articles from there stop appearing in results. Tags on a page (meta robots, X-Robots-Tag, Content-Usage) take effect the next time we read that page; to keep pages out before that, use robots.txt or the form.
  • Within 72 hours we delete the data we stored from your site, usually within minutes. Backups rotate within 7 days.
  • We keep a record of your choice (your domain, the date and the reason) so that we keep honoring it and, as proof that we did, a log of our requests to your site and the versions of your robots.txt, never the content of your articles, for up to two years. The contact details of whoever made the request are handled as our Privacy Policy says.
  • Our Terms also require our customers to stop using stored results from a publisher that opted out within 30 days of our notice.

Tell us what you need

Use this form for anything about your site: to leave, to show headlines and links only, to remove pages, to slow our crawler down or to talk about more. We never ask why, and we don’t argue.

  • You get an acknowledgment by email right away, with a case number.
  • Exclusions are applied within minutes; we commit to under 24 hours.
  • Stored data is deleted within 72 hours.
  • You get written confirmation when it’s done.
  • If the email you use is at your domain, you can confirm your request with one click. If it isn’t, we apply it anyway; for more than three sites, we check the rest first, within 24 hours.
What do you need?

We send the acknowledgment and your case number here.

For example, editor, product or legal.

One per line, like diarioejemplo.example.

Anything else we should know. You don’t have to tell us why.

Mentioned in the news? Choose “Remove results about me”. We review those requests case by case, as privacy law requires: see People in the news in our Privacy Policy.

Prefer email? Write to publishers@typesearch.ai.

How to verify our crawler

Every request says who it is and links to our crawler page. Without a browser, the user agents are:

Mozilla/5.0 (compatible; TypesearchBot/1.0; +https://typesearch.ai/bot)
Mozilla/5.0 (compatible; Typesearch-User/1.0; +https://typesearch.ai/bot)

Our crawler’s requests come from the addresses published at https://typesearch.ai/bot/ips.json: today, a single IPv4 address, 178.128.144.188. Each one is also signed with Web Bot Auth, and our public keys are at https://api.typesearch.ai/.well-known/http-message-signatures-directory. Reverse DNS verification isn’t available yet. If something uses our name from another address, or without our signature, it isn’t us: tell us at bot@typesearch.ai.

Licensing and partnerships

If you want more — your premium content under a license, longer excerpts, a partnership — let’s talk. Choose “Licensing or partnership” in the form, or write to publishers@typesearch.ai.

Contacts