Google Search rests on a social contract: their bots can crawl our sites, they can index our sites, and they can show excerpts of our sites because

korrupt@nrw.social

@RnDanger @inthehands well, we don’t know and we will see. My guess are separate scrapers (officially) and a lot of mistrust (are there others?) and masses of unidentified scrapers. Nevertheless, Google can better afford to play by the rules, since hey already own the largest index. Think also of Video etc. Will volume win the war? Or quality and freshness? Etc. Future is difficult.

schamschula@mastodon.social

@albertcardona @inthehands It involves a couple steps, given the idiosyncrasies of the nginx regex support (no full pcre here!).
I keep two classes of blocked agents: (1) bad agents; and (2) scrapping false agents. A third regex unblocks agents that are false positives (due to (2)).

gturri@climatejustice.social

@inthehands If I understand your question correctly (sorry if it's not the case) I think that Anubis, the AI crawler protection, could be part of the solution. Not only would that work for Google, that would (or at least *should*) also work against other crawlers.
Another advantage is that it can work along your other solutions.
OTHO the drawback is that it would work against all crawler, so you would "disappear" from every search engine...

https://en.wikipedia.org/wiki/Anubis_(software)