Fingers crossed Gnome follows suit! :)

you are viewing a single comment's thread
view the rest of the comments
[–] 7 points 3 days ago* (last edited 3 days ago) (3 children)

IBM Granite (EDIT: and Apertus) do disclose all their training data and claim that all their training data is effectively free of copyright (highly permissively licensed). I have not been able to verify that, due to my lack of skills with the conventions and tools of LLM / Agent training and publishing.

So, yeah, probably (EDIT: two one).

  • source
  • parent
  • hideshow 3 child comments
  • [–] [S] 8 points 3 days ago* (1 child)

    The Apertus Swiss AI unfortunately doesn't seem to live up to its claims, and they have been silent on the issue brought up there. I honestly suspect the same of IBM's Granite, but have not investigated.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 4 points 3 days ago* (last edited 3 days ago)

    Thank you for the link! It does look like Apertus itself might be Free Software (the U.S. copyright office says training can infringe, but is usually fair use), but it can still output derivative works of copyrighted inputs that might prevent them from being distributed as-is (for example, requiring attribution) -- at all, much less under a strong copyleft.

  • source
  • parent