Bill proposed to outlaw downloading Chinese AI models.

schizoidman@lemm.ee · 14 hours ago

Bill proposed to outlaw downloading Chinese AI models.

Gamers_mate@beehaw.org · 14 hours ago

I could understand banning closed source models but open sourced models that work better than anything propriety isn’t that just the free market that corporations like to pretend to be part of?

teawrecks@sopuli.xyz · 2 hours ago

It’s also the free market for those corporations to buy a government and use it to outlaw competition.

jarfil@beehaw.org · 12 hours ago

Define “open sourced model”.

The neural network is still a black box, with no source (training data) available to build it, not to mention few people have the alleged $5M needed to run the training even if the data was available.

thingsiplay@beehaw.org · 12 hours ago

Define “open sourced model”.

The term itself is actually shockingly simple. Source is the original material that was used to build this model, training data and all files that are needed to compile and create the model. It’s Open Source, if these files are available (preferably with an Open Source compatible license). It’s not. We only get binary data, the end result and some intermediate files to fine tune it.

jonne@infosec.pub · 12 hours ago

They were only for the free market if they could force it on others.

thingsiplay@beehaw.org · 13 hours ago

Well its still not Open Source.

Gamers_mate@beehaw.org · 12 hours ago

Is part of the code not available?

thingsiplay@beehaw.org · 12 hours ago

None of the code and training data is available. Its just the usual Huggingface thing, where some weights and parameters are available, nothing else. People repeat DeepSeek (and many other) Ai LLM models being open source, but they aren’t.

They even have a Github source code repository at https://github.com/deepseek-ai/DeepSeek-R1 , but its only an image and PDF file and links to download the model on Huggingface (plus optional weights and parameter files, to fine tune it). There is no source code, and no training data available. Also here is an interesting article talking about this issue: Liesenfeld, Andreas, and Mark Dingemanse. “Rethinking open source generative AI: open washing and the EU AI Act.” The 2024 ACM Conference on Fairness, Accountability, and Transparency. 2024

P03 Locke@lemmy.dbzer0.com · edit-2 5 hours ago

This literally took one click: https://github.com/deepseek-ai

Stop spreading FUD.

jarfil@beehaw.org · 5 hours ago

Where’s the training data?

P03 Locke@lemmy.dbzer0.com · 55 minutes ago

Nobody releases training data. It’s too large and varied. The best I’ve seen was the LAION-2B set that Stable Diffusion used, and that’s still just a big collection of links. Even that isn’t going to fit on a GitHub repo.

Besides, improving the model means using the model as a base and implementing new training data. Specialize, specialize, specialize.

jarfil@beehaw.org · 36 minutes ago

What about these? Dozens of TB here:

https://huggingface.co/HuggingFaceFW

There is also a LAION-5B now, and several other datasets.

Crotaro@beehaw.org · 4 hours ago

Does open sourcing require you to give out the training data? I thought it only means allowing access to the source code so that you could build it yourself and feed it your own training data.

jarfil@beehaw.org · 2 hours ago

Open source requires giving whatever digital information is necessary to build a binary.

In this case, the “binary” are the network weights, and “whatever is necessary” includes both training data, and training code.

DeepSeek is sharing:

NO training data
NO training code
instead, PDFs with a description of the process
binary weights (a few snapshots)
fine-tune code
inference code
evaluation code
integration code

In other words: a good amount of open source… with a huge binary blob in the middle.

teawrecks@sopuli.xyz · 2 hours ago

Is there any good LLM that fits this definition of open source, then? I thought the “training data” for good AI was always just: the entire internet, and they were all ethically dubious that way.

What is the concern with only having weights? It’s not abritrary code exectution, so there’s no security risk or lack of computing control that are the usual goals of open source in the first place.

To me the weights are less of a “blob” and more like an approximate solution to an NP-hard problem. Training is traversing the search space, and sharing a model is just saying “hey, this point looks useful, others should check it out”. But maybe that is a blob, since I don’t know how they got there.