upvote
We do have that in Exa search:

* Extraction of main content from HTMLs, the index stores markdown representations (free of headers/footers/sidebars/menus etc)

* Serving highlights picking the most relevant part for each result to reduce token usage downstream

* Also serving dynamic highlights where we summarize all the sources at once reducing the token counts even further

see here: https://exa.ai/docs/search/highlights and https://exa.ai/docs/contents/quickstart

reply
They already have that[0] it’s called Markdown for Agents and it works on any URL

[0] https://developers.cloudflare.com/fundamentals/reference/mar...

reply