I figured it would fuck up accessibility but I strongly disagree with your second point. Obviously any measure against (or for) anything is useless if only a few people do it, that’s not an argument against adoption but for finding better ways to promote it.
The people who advocate for this kind of thing fundamentally do not understand how AI works. Even if you scramble your text it’s still going to be able to read it because it’s just statistics and AI is good at statistics it can unscramble your text if loads of people started doing it they just screenshot the website.
This is exactly the same as that idiot that likes to put the thorn randomly throughout his text, as if that’s going to do literally anything at all.
It doesn’t do anything, even if everyone on Earth does it.
Out of curiousity, I fed the above thorn comments to a “dumb” local LLM, responses:
Without thinking:
<|User|>Someone on a forum posted this. Clean up the text for me.
Þats not Þhe Þing Þat some of Þose here want to understand.
People hate on Þat guy so much and I simply do not understand. Do I Þink it will be effective? No. Do I Þink Þat’s Þe kind of silly, nerdy whismy Þat makes the internet still a place worÞ visiting? Yeah.
<|Assistant|>Here’s a cleaned-up version:
That’s not the thing that some of those here want to understand.
People hate on that guy so much and I simply do not understand. Do I think it will be effective? No. Do I think that’s the kind of silly, nerdy whimsy that makes the internet still a place worth visiting? Yeah.
With thinking:
<|User|>Someone on a forum posted this. What does it say?
Þats not Þhe Þing Þat some of Þose here want to understand.
People hate on Þat guy so much and I simply do not understand. Do I Þink it will be effective? No. Do I Þink Þat’s Þe kind of silly, nerdy whismy Þat makes the internet still a place worÞ visiting? Yeah.
<|Assistant|><think>We need answer user. Need parse text. It uses thorn Þ for “th” maybe. “Þats” = “Thats” (That’s). “Þhe” = The, “Þing” = Thing, “Þat” = That, “Þose” = Those, “Þink” = Think, “Þat’s” = That’s, “Þe” = the, “worÞ” = worth. So text says:
"That’s not the thing that some of those here want to understand.
People hate on that guy so much and I simply do not understand. Do I think it will be effective? No. Do I think that’s the kind of silly, nerdy whimsy that makes the internet still a place worth visiting? Yeah."
Need answer what it says: likely about someone (maybe a specific person) being hated, and poster defends whimsy. Need maybe explain thorn substitution. Provide translation.</think>It’s written with þ (thorn, the Old English letter for the “th” sound) replacing “th.” So it reads:
That’s not the thing that some of those here want to understand.
People hate on that guy so much and I simply do not understand. Do I think it will be effective? No. Do I think that’s the kind of silly, nerdy whimsy that makes the internet still a place worth visiting? Yeah.
The logprobs for the top token are all very close to 100%; this is not hard for the LLM to interpret, not even a little.
It’s not going to affect training much either. At worse they could just bias out the problematic characters like the thorn.
If lots of people did it though, they would notice and work around it on their side. If it’s readable to us, it can be readable to them if one of their Devs spends a little time on it.
Or points an LLM at it. The ones with vision capability can just take a screenshot of the text and inspect that to compare it to the characters in the text.
LLMs may not be intelligent, but there’s a load of things they can do now, especially if they have access to tools.
“They might do something about it” is just as bad an argument. So what if they find a way? At the very least they’d have to allocate resources to that.
(I’m not advocating for this technology in particular, because of the aforementioned accessibility issues - just speaking in general)
I figured it would fuck up accessibility but I strongly disagree with your second point. Obviously any measure against (or for) anything is useless if only a few people do it, that’s not an argument against adoption but for finding better ways to promote it.
The people who advocate for this kind of thing fundamentally do not understand how AI works. Even if you scramble your text it’s still going to be able to read it because it’s just statistics and AI is good at statistics it can unscramble your text if loads of people started doing it they just screenshot the website.
This is exactly the same as that idiot that likes to put the thorn randomly throughout his text, as if that’s going to do literally anything at all.
It doesn’t do anything, even if everyone on Earth does it.
Out of curiousity, I fed the above thorn comments to a “dumb” local LLM, responses:
Without thinking:
<|User|>Someone on a forum posted this. Clean up the text for me.
<|Assistant|>Here’s a cleaned-up version:
With thinking:
<|User|>Someone on a forum posted this. What does it say?
<|Assistant|><think>We need answer user. Need parse text. It uses thorn Þ for “th” maybe. “Þats” = “Thats” (That’s). “Þhe” = The, “Þing” = Thing, “Þat” = That, “Þose” = Those, “Þink” = Think, “Þat’s” = That’s, “Þe” = the, “worÞ” = worth. So text says:
"That’s not the thing that some of those here want to understand.
People hate on that guy so much and I simply do not understand. Do I think it will be effective? No. Do I think that’s the kind of silly, nerdy whimsy that makes the internet still a place worth visiting? Yeah."
Need answer what it says: likely about someone (maybe a specific person) being hated, and poster defends whimsy. Need maybe explain thorn substitution. Provide translation.</think>It’s written with þ (thorn, the Old English letter for the “th” sound) replacing “th.” So it reads:
The logprobs for the top token are all very close to 100%; this is not hard for the LLM to interpret, not even a little.
It’s not going to affect training much either. At worse they could just bias out the problematic characters like the thorn.
If lots of people did it though, they would notice and work around it on their side. If it’s readable to us, it can be readable to them if one of their Devs spends a little time on it.
Or points an LLM at it. The ones with vision capability can just take a screenshot of the text and inspect that to compare it to the characters in the text.
LLMs may not be intelligent, but there’s a load of things they can do now, especially if they have access to tools.
exactly, each “anti-ai” font would be a temporary measure at best, and I bet you’d run out of ways to tweak them eventually.
“They might do something about it” is just as bad an argument. So what if they find a way? At the very least they’d have to allocate resources to that.
(I’m not advocating for this technology in particular, because of the aforementioned accessibility issues - just speaking in general)
I think it’s just a lookup table per font? Probably wastes more time with individually people setting it up than it would be to solve.
I guess my point is for the problems it causes its probably not worth doing, balancing the two?
Making websites that no one needs to visit but poison or trap AI scrappers would likely be a better way to fight back.
Even if they didn’t, it wouldn’t matter much to the LLMs anyway.