Hacker Newsnew | past | comments | ask | show | jobs | submit | eschaton's commentslogin

If you’re reading “Hacker” “News” either this sequence of English words should make sense to you or you should be competent enough to use a search engine and/or Wikipedia to easily learn what they mean.

Anyone who sells such an exploit should go to prison.

Agreed--if you can catch them of course.

The NetBSD developers keep a great CHANGES document up to date with what’s new overall, and even categorize it by architecture and machine when appropriate.

Want to know what changed between 10.1 and 11? Or 10.0 and 10.1? It’s all right there, summarized by the people who did the work.


Because simple similarity doesn’t usually result in a lawsuit? Especially if one of the words isn’t a public product name and also happens to be an existing word with no existing specific use in the industry.


Failing to defend trademarks for such similarity named and closely related products has extremely high risks.


“Xenon” is a pre-existing word. Intel’s not going to lose protection on “Xeon” for IBM using “Xenon” as an internal codename for their own CPU.


Pre-existing word like “Word” “Excel” “Apple” etc.

We know about “Xenon” because they used the internal name outside the company. There’s zero difference between an unofficial and official branding when you’re using it in press releases.


I didn’t see anything in there instructing the LLM not to generate text about goblins.


And they had also switched to NEXTSTEP on SPARC, HP-PA, and x86 hardware a couple years after the last NeXT hardware shipped. (It’s a trivial recompile for 99.9% of NeXT software.)


The Pyro accelerator.


1995


If someone were to train a coding-oriented LLM on only GPL code, I would assert that any code it outputs is covered by GPL because an LLM is fundamentally storing and reproducing its input, not understanding it and generating new things based on that understanding.


The starting point of GNU was that Stallman was pissed he got in trouble when he got caught copying code from the Symbolics sources to the MIT and LMI sources, which was against the agreement Symbolics and LMI had with the AI Lab, which was that improvements could only flow one-way (AI Lab to commercial). Dan Weinreb (RIP) confirmed this publicly.

Of course, not long after starting GNU, Stallman got caught copying code from Unipress emacs sources into the then-new GNU emacs sources. Oops! That’s why it was difficult for quite a long time to find early GNU emacs sources online—they were purged from various archives because they were infringing.


I have personally seen this happen:

Someone tried to contribute “vibe-coded” device support to a project I’m involved with, they said they did it all based on the device documentation, the code their agents spit out was copied verbatim out of a (GPL’d) project with which I’m familiar which supports that device.

LLMs are not learning things and then using that learning to construct new things. They are essentially a form of lossy compression of their training set. And you don’t need to be explicit about trying to reproduce a portion of that training set for an LLM to output one.


As it happens, all evaluations I have seen in the news were in fact explicit about trying to reproduce a portion of the training set.

I am not aware of any study attempting to measure unintentional reproduction.

With your example, I question whether you have seen this happen first hand. For all I know, the contributor could have explicitly prompted the model to reference the GPL project and had the agent clone the code from the web.


At a certain point you have to take people at their word; I’m reporting what the contributor said they did (used the documentation to generate the code).


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: