• 0 Posts
  • 10 Comments
Joined 3 years ago
cake
Cake day: January 26th, 2024

help-circle

  • The IP law is purely a function of economic power, so we know how that will go. To the maximum benefit of the plutocrats that rule our global civilization.

    But the more interesting part is the philosophical discovery: If a piece of code can be generated from scratch by a machine that is not creative, then any comparable piece of code is also not creative. In other words, only the very first time a new type of code or algorithm is written can there be creativity - if one follows your argument that LLM machines cannot be creative. This essentially raises the bar for copyright, and makes an objective test of what is copyrightable or not possible. If you can describe something in a prompt, and get a result of comparable function and quality to a human written version, that code is not creative either, therefor not copyrightable.

    Obviously we’re not quite there yet, but the more LLMs improve, the less code becomes copyrightable.

    Coding becomes pure craftsmanship, like cutting and sanding some lumber and nailing it together to build a shed.

    So if we were a rational and logical species, the overwhelming amount of copyright would be deleted and become public domain. Because it is now worth only the cheap solar power needed to generate it.



  • Not true, this does not occur frequently. This study and software LiCoEval from 2024 found 0.88% to 2.01% of code “strikingly similar to existing open-source implementations”. Afaik this is mostly textbook examples, snippets from stack overflow snippets or common github repositories, often replicated api examples and language boilerplate. How you prompt and refine also matters, and for generating novel code or business logic the LLM simply cannot use memorized snippets.

    Presumably since then LLMs have worked to reduce that number of memorized code. Since LLMs cannot memorize all their training data, that number is limited. LiCoEval can find the often memorized examples and train to remove them, or suppress them, or they find other ways to reduce direct reproduction from memorization. For example it would be possible to do what malus.sh does with all the training data. Then it cannot memorize copyrighted code.

    So for a model that came out 2026 this already small number might not be that relevant anymore.




  • Well they will probably be massively over capacity. But then you could run the for half a day during sunlight hours and just use solar. Solar PM and wind kite power basically gives us near infinite super cheap energy at certain times, so you could still make use of those compute centers. Turn them off during the night, modernize the cooling and noise dampening, add some solar and you can use the compute to run models to create new medicines like anti-cancer meds or protein folding. Unlikely that’s going to happen but that is what a sane civilization would do with all that compute.


  • “internet indexer” should really be a government / EU / UN sponsored public utility thing. And then you have front ends that can improve the search query and filter and sort the results.

    I don’t know how much it would cost / how large a data center it would take to cache the web pages like google does. But there is just no reason to have so much wealth / advertising drain from the public for such a fundamental thing.