▲ 424 ▼ Cara is finally taking action against AI-bro scrapers (europe.pub) submitted 1 month ago* (last edited 1 month ago) by destructdisc@lemmy.world to c/fuck_ai@lemmy.world 53 comments fedilink hide all child comments Context: https://blog.cara.app/blog/cara-needs-your-help
[–] AeonFelis@lemmy.world 4 points 1 month ago (3 children) Don't they use Anubis? permalink fedilink source hideshow 3 child comments replies: [–] ruby@lemmy.dbzer0.com 3 points 1 month ago (2 children) i just visited and no, they don't seem to use it, or at least i could access the site without it. but even if they did, anubis is incapable of defending against targeted scraping. permalink fedilink source parent hideshow 2 child comments replies: [–] emmowo@lemmy.world 1 point 1 month ago* (1 child) Anubis can do an okay job... (at least so the whole site isn't being actively DDos'ed) at the cost of compromising user experience a lot by sending harder challenges more often. It kinda sucks that the average user might not have that fast of a CPU, while motivated scrapers might spend tonnes on compute just out of spite. (i am a bit biased on this though) permalink fedilink source parent hideshow 1 child comment replies: [–] ruby@lemmy.dbzer0.com 2 points 1 month ago anubis does its thing by making you waste your cpu cycles and then giving you a cookie, if you use that cookie for subsequent requests then you're free to browse the site. it might in theory slow down dumb scrapers that don't save that state and parallelize the mass scraping via multiple ip addresses (since without the cookie they'll have to solve the challenge for each request) but if one's making a targeted attack for one specific site, they don't do any of that. it should be trivial for a bad actor to solve the challenge once and use the cookie, like a regular user would. anubis only slows down scrapers that scrape all sites indiscriminately and don't expect it, once you know it's there then it's easy to bypass (and if you can afford to scrape the entire site you can surely afford to get past anubis once). permalink fedilink source parent
[–] ruby@lemmy.dbzer0.com 3 points 1 month ago (2 children) i just visited and no, they don't seem to use it, or at least i could access the site without it. but even if they did, anubis is incapable of defending against targeted scraping. permalink fedilink source parent hideshow 2 child comments replies: [–] emmowo@lemmy.world 1 point 1 month ago* (1 child) Anubis can do an okay job... (at least so the whole site isn't being actively DDos'ed) at the cost of compromising user experience a lot by sending harder challenges more often. It kinda sucks that the average user might not have that fast of a CPU, while motivated scrapers might spend tonnes on compute just out of spite. (i am a bit biased on this though) permalink fedilink source parent hideshow 1 child comment replies: [–] ruby@lemmy.dbzer0.com 2 points 1 month ago anubis does its thing by making you waste your cpu cycles and then giving you a cookie, if you use that cookie for subsequent requests then you're free to browse the site. it might in theory slow down dumb scrapers that don't save that state and parallelize the mass scraping via multiple ip addresses (since without the cookie they'll have to solve the challenge for each request) but if one's making a targeted attack for one specific site, they don't do any of that. it should be trivial for a bad actor to solve the challenge once and use the cookie, like a regular user would. anubis only slows down scrapers that scrape all sites indiscriminately and don't expect it, once you know it's there then it's easy to bypass (and if you can afford to scrape the entire site you can surely afford to get past anubis once). permalink fedilink source parent
[–] emmowo@lemmy.world 1 point 1 month ago* (1 child) Anubis can do an okay job... (at least so the whole site isn't being actively DDos'ed) at the cost of compromising user experience a lot by sending harder challenges more often. It kinda sucks that the average user might not have that fast of a CPU, while motivated scrapers might spend tonnes on compute just out of spite. (i am a bit biased on this though) permalink fedilink source parent hideshow 1 child comment replies: [–] ruby@lemmy.dbzer0.com 2 points 1 month ago anubis does its thing by making you waste your cpu cycles and then giving you a cookie, if you use that cookie for subsequent requests then you're free to browse the site. it might in theory slow down dumb scrapers that don't save that state and parallelize the mass scraping via multiple ip addresses (since without the cookie they'll have to solve the challenge for each request) but if one's making a targeted attack for one specific site, they don't do any of that. it should be trivial for a bad actor to solve the challenge once and use the cookie, like a regular user would. anubis only slows down scrapers that scrape all sites indiscriminately and don't expect it, once you know it's there then it's easy to bypass (and if you can afford to scrape the entire site you can surely afford to get past anubis once). permalink fedilink source parent
[–] ruby@lemmy.dbzer0.com 2 points 1 month ago anubis does its thing by making you waste your cpu cycles and then giving you a cookie, if you use that cookie for subsequent requests then you're free to browse the site. it might in theory slow down dumb scrapers that don't save that state and parallelize the mass scraping via multiple ip addresses (since without the cookie they'll have to solve the challenge for each request) but if one's making a targeted attack for one specific site, they don't do any of that. it should be trivial for a bad actor to solve the challenge once and use the cookie, like a regular user would. anubis only slows down scrapers that scrape all sites indiscriminately and don't expect it, once you know it's there then it's easy to bypass (and if you can afford to scrape the entire site you can surely afford to get past anubis once). permalink fedilink source parent