This story should terrify any open source project or company that allows LLM-generated code into their project.
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
"And that’s something that requires human judgement."
someone in microsoft is coming up with a paid extra "ai compliance centre" which will aim to do this for you.
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
@david_chisnall s/Gemini/Claude/?
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
@david_chisnall critical support for AI companies totally destroying copyright law as one of it's casualties. Hopefully we can count that as a positive outcome of all this while eating our wacky cake and water pies after the crash
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
@david_chisnall It’s absolutely an issue, in this case I’d be wary of the source, the author doesn’t have the best track record as far as being thruthful about this app goes: https://daringfireball.net/2026/08/retraction_app_store_rejection_of_the_week#:~:text=frustrated%20by%20the,the%20domain%20darkhours.app.
-
@MaddieM4 @david_chisnall isnt this the app that got rejected from the App Store for being an astrology app passing off as an astronomy app?
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
@david_chisnall@infosec.exchange
Well, ultimately, it's how the #PayToTrain #BusinessModel works.
The original "author" #vibecoded the project using the very same #agent (#claude).
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
@david_chisnall jesus fucking christ
-
@david_chisnall critical support for AI companies totally destroying copyright law as one of it's casualties. Hopefully we can count that as a positive outcome of all this while eating our wacky cake and water pies after the crash
@fluffykittycat @david_chisnall The idea that you're going to get a destruction of copyright that applies equally to *you* as to VC-backed monsters is a delusional fantasy.
Even if you did get this in some jurisdictions, you would not get it uniformly across all jurisdictions.
Free software needs to be free regardless of whose version or extent of copyright abolution a user might be subject to the jurisdiction of.
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
@david_chisnall I should note that the inadvertently plagiarized web app is *also* coded with Claude. Who knows what parts of previous projects it incorporates?
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
@david_chisnall dang double bummer that that person is a sloperator. some of his stuff seemed fun in passing and i never really checked into it
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
This is a critical "detail", seemingly obvious once you've made it clear like this, that i suspect small hobbiest type users are unaware of.
I assume large corporate users know this already and are ok eith just stealing copyrighted works and hiding behind lawyers and legal costs.
Thsnk you for this.
-
@fluffykittycat @david_chisnall The idea that you're going to get a destruction of copyright that applies equally to *you* as to VC-backed monsters is a delusional fantasy.
Even if you did get this in some jurisdictions, you would not get it uniformly across all jurisdictions.
Free software needs to be free regardless of whose version or extent of copyright abolution a user might be subject to the jurisdiction of.
@dalias @david_chisnall you're right, but the damage to the legitimacy of copyright ideology is still happening.
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
@david_chisnall The "Mea Culpa" author's *mea culpa* doesn't pass the smell test.
The first sign of trouble is the identical name for a nearly identical app.
When I wrote a small astronomy calculation library (under a name that would be familiar to astronomers, so likely not unique) I made *damn sure* that there was nothing else out there on Github or Codeberg with the same name that was sufficiently like it to incur possible trademark issues.
-
@david_chisnall The "Mea Culpa" author's *mea culpa* doesn't pass the smell test.
The first sign of trouble is the identical name for a nearly identical app.
When I wrote a small astronomy calculation library (under a name that would be familiar to astronomers, so likely not unique) I made *damn sure* that there was nothing else out there on Github or Codeberg with the same name that was sufficiently like it to incur possible trademark issues.
@david_chisnall Also, as a longstanding amateur astronomer, I don't particularly trust the original app (https://darkhours.app) that he almost certainly plagiarized. I plugged in the coordinates for a location I did a little light observing from last month, which it claimed was "Bortle 3" (very dark sky, minimally affected by city lights). That's broadly consistent with what I saw myself -- some light pollution domes from cities in the distance, but very dark overhead. Moving just a little over *one hundred meters* to the north changed it to "Bortle 6" (suburban levels of light pollution; very much visibly lighter than Bortle 3). That's ... not how light pollution works, *at all*.
-
@david_chisnall It’s absolutely an issue, in this case I’d be wary of the source, the author doesn’t have the best track record as far as being thruthful about this app goes: https://daringfireball.net/2026/08/retraction_app_store_rejection_of_the_week#:~:text=frustrated%20by%20the,the%20domain%20darkhours.app.
@joostvanderborg Came here to post the DF article but I see you posted it already. I agree that Godier’s credibility is questionable.
Maybe that changes some of the facts behind @david_chisnall’s post but the issues it raises around provenance and copyright of LLM remain, regardless of Godier’s intent in this particular incident. The question I have is whether Godier actually used Claude or whether he himself made the copy and is trying to blame Claude.
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
@david_chisnall @zzt honestly this story will not terrify anywhere near as many people as it should
-
This story should terrify any open source project or company that allows LLM-generated code into their project.
Someone created a new app, using Claude. Only it wasn’t a new app, it was clearly plagiarised because it happened that there was an example in the training set that exactly matched the requirements. The similarity was well within the range that courts have previously used to determine a derived work.
When you have this level of similarity, the requirement comes to you to prove that there was no way that the original work could have flowed to your project. Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
Just because the US copyright office has ruled that you, as the person promoting an LLM, cannot assert copyright on the output, does not mean that someone else can’t. If you take a DVD and transcode it to H.264, there is no creative step and so the new copy is not something subject to independent copyright, but it is a derived work of the DVD (itself a lower-quality derived work of the original masters) and so subject to the same copyright.
Importantly in this story, the person prompting Claude had no idea that the original app existed. To safely use the code, they would need to search everything in the training data and discard outputs that would meet the bar of being substantially similar. And that’s something that requires human judgement.
@david_chisnall local LLM's are the new printing press for a house. Major companies only have a monopoly bc there's an illusion they are the only ones capable
Get your opensource model today guys!
-
@david_chisnall Also, as a longstanding amateur astronomer, I don't particularly trust the original app (https://darkhours.app) that he almost certainly plagiarized. I plugged in the coordinates for a location I did a little light observing from last month, which it claimed was "Bortle 3" (very dark sky, minimally affected by city lights). That's broadly consistent with what I saw myself -- some light pollution domes from cities in the distance, but very dark overhead. Moving just a little over *one hundred meters* to the north changed it to "Bortle 6" (suburban levels of light pollution; very much visibly lighter than Bortle 3). That's ... not how light pollution works, *at all*.
@dpnash @david_chisnall from quick peek at github this app is also fully vibe coded slop, so I wouldn't be surprised that it's full of shit.
Also there's a non-zero chance that the app that's supposedly plagiarised is a plagiarism in of itself, and they simply both put in the same prompts and got same output..
It's slop all the way down.
-
@dpnash @david_chisnall from quick peek at github this app is also fully vibe coded slop, so I wouldn't be surprised that it's full of shit.
Also there's a non-zero chance that the app that's supposedly plagiarised is a plagiarism in of itself, and they simply both put in the same prompts and got same output..
It's slop all the way down.
@siwek@tooinconsistent.com @david_chisnall As far as I can tell, the app "bins" light pollution estimates every 0.01 degree of latitude (and possibly longitude as well, but I didn't check that variable), so there are discontinuities when that decimal rolls over. That wouldn't be too bad if the app's model for calculating light pollution levels was anything like accurate. An accurate model wouldn't jump suddently from Bortle 3 to Bortle 6 over a distance like that (0.01 degrees of latitude is only about a km) -- there's nowhere on Earth where light pollution levels change that abruptly.
This fairly reputable source (lightpollutionmap.info) has Bortle 4 across the entire area for miles around where I was looking, with a slight brightening (still Bortle 4) in a small town and a slight darkening (close to Bortle 3 in its darkest spots) outside of town, and is *far* closer to the reality of the situation: https://www.lightpollutionmap.info/#zoom=10.18&lat=48.5140&lon=-123.0300&state=eyJiYXNlbWFwIjoiTGF5ZXJPU00iLCJvdmVybGF5Ijoic2JfMjAyNSIsIm92ZXJsYXljb2xvciI6ZmFsc2UsIm92ZXJsYXlvcGFjaXR5IjoiNjAiLCJmZWF0dXJlc29wYWNpdHkiOiI4NSJ9