> how do you prove a human wrote something? Forget about the why or the value in it, just: how?
> Semoi is a plugin (currently only available for Obsidian) which tracks the length of time it took for a document to be written up
Trying to mechanistically prove that a human created some content as opposed to ai, in the age of LLMs and style transfer when you can just ask for something to be written in the style of Mark Twain or drawn in a style of van Gogh and get a great output, is a fool's errand.
All solutions to this end are going to be some form of attestation.
Even the proposed approach of tracking keystrokes and timing as a form of mechanical attestation, is going to be short lived because someone will train an AI on a corpus of human keystrokes and get a replication. May not even need an ai for this, a stochastic program could conceivably reproduce this behavior.
> “I didn’t think anyone would care” prevented me from writing, though
A few years ago I started writing Twitter threads [0]. A few weeks ago I passed 200 total threads.
When I started writing them, my thought process was: "Is anyone going to be interested in my stories/ideas??"
Dear HN comment reader, I can 100% assure you of two things:
1. If you write things, at least one person will read them.
2. It is VERY hard to predict what people will find interesting
e.g. some of the threads I thought people would find the least interesting got the most traction and vice versa. The only way to find out is to write it down.
I would also add that just writing, a LOT, helps you become a better thinker and writer. Twitter threads in particular are great as they force you to distill a story down into bite sized chunks.
One additional benefit: you meet amazing people when you write about what you are interested in. Why? Because if someone likes your writing, they would probably like talking to you and you to them.
I recall a similar quote from Elton John that I'll paraphrase:
"I've had a lot of hits so you'd think I'd know in advance which ones will become hits. Songs that I was sure would become hits went nowhere and some songs that I didn't think anything of became my biggest hits"
I think we are reaching a point where some kind of solid attestation that things are not AI generated would be very valuable, for all kinds of different media (printed word, photo, video, etc).
I'm not sure how you would actually do any of that attestation, if it's even possible. Text seems especially difficult. Maybe photo/video could be achieved with specialized hardware and cryptographic signing though I don't know much about either so I'm not sure how it would work. Maybe all that attestation would also tie in to some kind of universal personal identifier online, so that bad actors can be tracked or excluded and can't repeatedly spin up new accounts.
It might mean a big reduction in privacy for certain online spaces that opt in to such a system... but the alternative of all trust being eroded and voices drowned out by a sea of bots or generated content seems potentially worse.
I have a slightly different approach on my blog https://ezeugo.dev where my entire thought process (edits, original ideas, rewrites are a part of the actual essay). By making the event stream part of the product, it shows the process and output as one.
Detecting Ai use based on "how many keystrokes did you type and when" doesn't solve the problem, because someone can just write a program to mimic human typing in the words.
So the issue isn't "did a human write something", it's what the actual content is
i do actually think the general thrust of the idea is good. i also think it's sort of inherently cursed, and has some amount of a predeterminable fate.
if it doesnt take off, it dies.
if it does take off and becomes a relevant currency of some sort, it will need improving. if it needs improving how far are we willing to go? does some alternative system fork off to handle severity of provenance concern?
what happens when AI action becomes indistinguishable from human action? what happens when sticking computers in your head becomes vogue?
if it takes off and fills a small niche, maybe thats the best future.
Like I said, cool idea, but seemingly very cursed from the get go.
There's plenty of microbehavioral analysis we can do that is initially effective but will get bypassed (with GANs being the purest way, or something more domain-specific).
You could imagine a livestream that's permanently published somewhere. But the verification of the livestream takes longer than reading the piece itself (and is itself vulnerable to faking).
Personally I think it comes back to something like writing under your real name - staking your reputation - as the most trustworthy indicator. At least people in your circles can trust you.
You can read ahead as to how this would go by looking at the CAPTCHA world, which unbeknownst to a lot of people left behind "click this image" as the actual test a long time ago and does a lot of behavioral analysis of mouse motions and stuff. Which is itself a constant arms race. Which I would still characterize as "advantage attacker" with regard to that aspect of it, CAPTCHAs still have at least some utility more from the other streams they have access to that are harder to fake, e.g., "this IP address is from a rentable service space and thus more likely to be a bot" sorts of checks.
That means this basically works, as long as it never gets large enough to attack. Which may suit the author just fine. Not everything has to solve the world's problems. But it won't generalize very far, no.
As others have pointed out, it's relatively a lot of effort to create an artifact that realistically current systems can pretty well forge.
I don't know that there is a scalable and comfortable solution to this problem (or at least one that is scalable and comfortable proportional to the demand for it).
Just like we trained on human language to create LLMs, we can train on human keystrokes with a similar algorithm and spit out believable (at least statistically) "human" keystrokes generated by machine.
Yeah I think the "how can you prove text was handwritten" question is a subset of the larger "how can you prove that a computer is being driven by a human" problem that all of the work around captcha, attestations, biometrics, and government-id auth has been aimed at. The fundamental issue, it seems to me, is that any signal that a human can provide to a computer (keystroke, camera frame, mouse click, etc) is inherently only parsable by code because a sensor has translated the analog signal into a digital one. That same requirement also ensures that the input can be digitally spoofed or automated. There's a similar problem on the output side: how can an analog user trust a digital certificate? What's stopping me from copying the certificate HTML or taking a screenshot and using it to trick people into thinking my AI content is handwritten?
I don't have any suggestions. I worry that the only strong solutions require a lot of power to be given to a centralized authority.
> One could work around Semoi by, for example, typing out a bunch of gibberish, leaving their editor open, and then pasting in an LLM generated texting and minting the proof. To which I would respond: why? That’s really pathetic.
Even before LLMs, proving that a specific human wrote some piece of text was difficult or imprecise due to coauthoring, editing, plagiarism, etc. The important question then, as now, is rather to determine whether a specific human approved some piece of text for publication under their name.
I'll be honest, I think this is a social issue that you're trying to apply a process-solution to solve.
If someone writes "this was completely hand-written with no AI assistance", I'll just believe them. I'm already committed to letting your words fill my brain for a bit, so I don't know why I would NEED a cryptographic signature to PROVE you aren't lying about WHO wrote it.
Being called out as a lier will be a LOT more painful than 1) using an LLM to write quickly and not lying, or 2) doing what you claim to be doing and writing it with your feeble, non-metallic human hands.
(First version of this comment had an example from an HN thread of this SPECIFIC behaviour getting called out but ehh that’s not the right vibe. My point is it does happen.)
That's actually a good point. Plenty of people just want to let their beliefs out in public. And there's not much benefit to lie about AI use. Some people use it and some people don't, and it's two very different groups of people.
I used to live with a ghostwriter for Simon Sinek, she was never attributed or acknowledged in "his" books. Since Ai authorship has become a common topic, I've wondered how the two means of producing a book relate and how it might inform this debate.
I mean the ability to fake all these metrics is relatively trivial, since I have written anti-bot fingerprinting scripts for automating various online services that want to keep you from automating them, you have the text you want to send from source A, you have random typing speed array you want to type them in, your chance of mistakes (put wrong character, backspace to remove, put in correct character).
And this doesn't even have to deal with all the stuff about mouse movements that you don't register?
Of course maybe I am just being typical programmer here, I guess lots of the people use generative AI would be defeated by copying pasting in the text and getting labeled AI, but that would also incorrectly label lots of people who have old texts in handwritten form they do not want to type all over again (of which I am one), and finally I assume that there is money in the field so producing something that allows bots to display "human heuristics" would probably get made and be profitable.
Funnily enough when I was automating things, generally twitter, I discovered that my real usage often got registered as bot, so I figured what's the use.
Also talking with someone who actually worked on bot-recognition by usage metrics said I was overly paranoid on some of the things I made my scripts do to appear human.
on the other hand - my automation was based on not wanting to spam services but provide the minimum level of content posting and regularity to benefit from algorithms that boost content based on the poster's engagement level, automation that wants to spam cannot benefit from this because slowing things down to show as if it was made in real time by a human (with fully non-headless browser etc.) does somewhat defeat the ability to spam like a machine.
I’ve been playing around with some similar ideas in a couple of pet projects. I truly believe that some type of “proof of work” system like this will be the only reliable way to have confidence in human writing. The various AI “detector” software that’s out there seems to be a dead end to me.
I'm choosing to read this as an option for people who want to put forward a serious attestation that the writing is not generated by AI, and are looking for ways to make that pledge tangible and give it a level of authority that feels more significant than "I promise" by getting that event-stream sent and signed. In that light, it seems like a useful offering!
I think it's important to consider the risks/rewards/benefits. There's definitely a sense in my circle of contacts that if a work is seen as purely human then it's somehow better and more authentic, and annecdotally, it's also possible to win points by taking something created with the help of AI and passing it off as your own independent work. Like social credit, there's a sense that you'll seem smarter than you feel yourself to be.
With that in mind, absolutely any technical solution to detect AI or attest to human authorship will be abused. The only context that a solution like this one supports is one where there isn't a risk or reward, the author just sincerely wants you to know that it's human-produced.
Plenty of people, after all, produce a draft with AI, then type it all out again editing and refining and updating as they go, and then ask an AI to look at the result and make suggestions, and then go back and make the changes they agree with. Such a workflow would be deemed human by semoi - so it's lucky that there's no point in lying about it.
This is a good idea - and I might pick it up for my writing, since a lot of it still happens in Obsidian. If a document is worked on across multiple days / revisions / app sessions - how is that handled?
i dunno why we dont do what the art world does and just blacklist you from everything if you're caught stealing/tracing and even revoke your degree(s) for it.
ai detectors always have so many false positives that im convinced they're only put in place by people too dumb to know any better and snakeoil salesmen.
im sorry but have you met the obsessive llm weirdos like
theyre already like that.
also im not talking about llm output? just... if you're caught passing it as your own in any field you get blacklisted which we already do with regular stealing
> One could work around Semoi by, for example, typing out a bunch of gibberish, leaving their editor open, and then pasting in an LLM generated texting and minting the proof. To which I would respond: why? That’s really pathetic.
I mean, I also think it's really pathetic to have an AI write something and then say "this text was written with no AI assistance"†. So if we've acknowledged that we're only going to stop non-pathetic people, why not skip the cryptographic hash signing and just go with the no-AI statement?
I understand that the goal is only to make lying hard, not impossible. However, I don't think this solution makes lying harder enough to meaningful.
As a writer who publishes regularly on the web, I just keep the edit history of my longer articles on my github. Its not perfect and you can never really prove you wrote something unless you do it in front of an observer watching you write in real life imo. At some point though, you have to make a choice between not wanting to waste people's and your readers' time but also privacy and additional effort. One of my recent blogs: https://decodingvibes.com/blog/what-we-talk-about-when-we-ta...
The technical solution is interesting but I think the only way to truly go about it is web of trust. Everyone knows that my work doesn't use AI because I hate it so much and so many people know that I don't use AI. It's embedded into the core of my personality. These tools can always be gamed but a true belief against AI cannot.
A lot of people do, that's the point. And over time there will be zero evidence that I ever used an LLM because I haven't. Many people in real life also believe it and they can also vouch for me.
"No information about your text is ever sent to the server, and there’s no way to identify the author based on the minted certificate"
This means the certificate is independent of author and source text.
There's nothing stopping you from sending fake counts/duration to the semoi server. It's a certificate that only says "at this point in time, this is the information I was provided with".
You can then attach it to any piece of text you like.
At the very least, you'd need the ability to prove that there is an underlying event stream with these characteristics, and that this exact event stream creates the document in question. You still can fake that event stream, but it becomes enough work to distract at least casual abusers.
But really, it's the equivalent of saying "I wrote this without AI, honest" in-doc and signing that with your personal key. The value depends entirely on your willingess to be truthful. (IOW: I predict we'll see a resurgence of reputation systems, to some extent)
I'm not doing that. I'm not going to be coerced into a writing style that's not mine by LLMs. AI detectors think that 30% of the stuff I wrote 20 years ago is AI-generated. If someone wants to think, incorrectly, that my writing was done by an LLM, I can't stop them.
> Semoi is a plugin (currently only available for Obsidian) which tracks the length of time it took for a document to be written up
Trying to mechanistically prove that a human created some content as opposed to ai, in the age of LLMs and style transfer when you can just ask for something to be written in the style of Mark Twain or drawn in a style of van Gogh and get a great output, is a fool's errand.
All solutions to this end are going to be some form of attestation.
Even the proposed approach of tracking keystrokes and timing as a form of mechanical attestation, is going to be short lived because someone will train an AI on a corpus of human keystrokes and get a replication. May not even need an ai for this, a stochastic program could conceivably reproduce this behavior.
A few years ago I started writing Twitter threads [0]. A few weeks ago I passed 200 total threads.
When I started writing them, my thought process was: "Is anyone going to be interested in my stories/ideas??"
Dear HN comment reader, I can 100% assure you of two things:
1. If you write things, at least one person will read them.
2. It is VERY hard to predict what people will find interesting
e.g. some of the threads I thought people would find the least interesting got the most traction and vice versa. The only way to find out is to write it down.
I would also add that just writing, a LOT, helps you become a better thinker and writer. Twitter threads in particular are great as they force you to distill a story down into bite sized chunks.
One additional benefit: you meet amazing people when you write about what you are interested in. Why? Because if someone likes your writing, they would probably like talking to you and you to them.
0 - https://x.com/alexpotato/status/2012723178577985948?s=20
"I've had a lot of hits so you'd think I'd know in advance which ones will become hits. Songs that I was sure would become hits went nowhere and some songs that I didn't think anything of became my biggest hits"
It's been a long time since I heard this so I'm probably mangling it. A quick search shows that he probably did say something like this though https://www.birminghammail.co.uk/news/showbiz-tv/sir-elton-j...
The lesson is - just put it out there and see what happens
I'm not sure how you would actually do any of that attestation, if it's even possible. Text seems especially difficult. Maybe photo/video could be achieved with specialized hardware and cryptographic signing though I don't know much about either so I'm not sure how it would work. Maybe all that attestation would also tie in to some kind of universal personal identifier online, so that bad actors can be tracked or excluded and can't repeatedly spin up new accounts.
It might mean a big reduction in privacy for certain online spaces that opt in to such a system... but the alternative of all trust being eroded and voices drowned out by a sea of bots or generated content seems potentially worse.
So the issue isn't "did a human write something", it's what the actual content is
if it doesnt take off, it dies.
if it does take off and becomes a relevant currency of some sort, it will need improving. if it needs improving how far are we willing to go? does some alternative system fork off to handle severity of provenance concern?
what happens when AI action becomes indistinguishable from human action? what happens when sticking computers in your head becomes vogue?
if it takes off and fills a small niche, maybe thats the best future.
Like I said, cool idea, but seemingly very cursed from the get go.
There's plenty of microbehavioral analysis we can do that is initially effective but will get bypassed (with GANs being the purest way, or something more domain-specific).
You could imagine a livestream that's permanently published somewhere. But the verification of the livestream takes longer than reading the piece itself (and is itself vulnerable to faking).
Personally I think it comes back to something like writing under your real name - staking your reputation - as the most trustworthy indicator. At least people in your circles can trust you.
That means this basically works, as long as it never gets large enough to attack. Which may suit the author just fine. Not everything has to solve the world's problems. But it won't generalize very far, no.
https://github.com/humthentic/itypedmypaper-v1
As others have pointed out, it's relatively a lot of effort to create an artifact that realistically current systems can pretty well forge.
I don't know that there is a scalable and comfortable solution to this problem (or at least one that is scalable and comfortable proportional to the demand for it).
I don't have any suggestions. I worry that the only strong solutions require a lot of power to be given to a centralized authority.
If someone writes "this was completely hand-written with no AI assistance", I'll just believe them. I'm already committed to letting your words fill my brain for a bit, so I don't know why I would NEED a cryptographic signature to PROVE you aren't lying about WHO wrote it.
Being called out as a lier will be a LOT more painful than 1) using an LLM to write quickly and not lying, or 2) doing what you claim to be doing and writing it with your feeble, non-metallic human hands.
(First version of this comment had an example from an HN thread of this SPECIFIC behaviour getting called out but ehh that’s not the right vibe. My point is it does happen.)
And this doesn't even have to deal with all the stuff about mouse movements that you don't register?
Of course maybe I am just being typical programmer here, I guess lots of the people use generative AI would be defeated by copying pasting in the text and getting labeled AI, but that would also incorrectly label lots of people who have old texts in handwritten form they do not want to type all over again (of which I am one), and finally I assume that there is money in the field so producing something that allows bots to display "human heuristics" would probably get made and be profitable.
Funnily enough when I was automating things, generally twitter, I discovered that my real usage often got registered as bot, so I figured what's the use.
Also talking with someone who actually worked on bot-recognition by usage metrics said I was overly paranoid on some of the things I made my scripts do to appear human.
I think it's important to consider the risks/rewards/benefits. There's definitely a sense in my circle of contacts that if a work is seen as purely human then it's somehow better and more authentic, and annecdotally, it's also possible to win points by taking something created with the help of AI and passing it off as your own independent work. Like social credit, there's a sense that you'll seem smarter than you feel yourself to be.
With that in mind, absolutely any technical solution to detect AI or attest to human authorship will be abused. The only context that a solution like this one supports is one where there isn't a risk or reward, the author just sincerely wants you to know that it's human-produced.
Plenty of people, after all, produce a draft with AI, then type it all out again editing and refining and updating as they go, and then ask an AI to look at the result and make suggestions, and then go back and make the changes they agree with. Such a workflow would be deemed human by semoi - so it's lucky that there's no point in lying about it.
ai detectors always have so many false positives that im convinced they're only put in place by people too dumb to know any better and snakeoil salesmen.
also im not talking about llm output? just... if you're caught passing it as your own in any field you get blacklisted which we already do with regular stealing
wtf is even your reply my dude.
I mean, I also think it's really pathetic to have an AI write something and then say "this text was written with no AI assistance"†. So if we've acknowledged that we're only going to stop non-pathetic people, why not skip the cryptographic hash signing and just go with the no-AI statement?
I understand that the goal is only to make lying hard, not impossible. However, I don't think this solution makes lying harder enough to meaningful.
We believe you on internet karma?
This means the certificate is independent of author and source text.
There's nothing stopping you from sending fake counts/duration to the semoi server. It's a certificate that only says "at this point in time, this is the information I was provided with".
You can then attach it to any piece of text you like.
At the very least, you'd need the ability to prove that there is an underlying event stream with these characteristics, and that this exact event stream creates the document in question. You still can fake that event stream, but it becomes enough work to distract at least casual abusers.
But really, it's the equivalent of saying "I wrote this without AI, honest" in-doc and signing that with your personal key. The value depends entirely on your willingess to be truthful. (IOW: I predict we'll see a resurgence of reputation systems, to some extent)
It's so difficult to write, you review, edit, edit and you read it and you can't unsee the LLM criticism.