Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Any chance of this being released open weights? Or is the risk of bad PR too high (especially near a US election)?

It being 30B gives me hope.



Meta text to image model cm3leon[0], was announced july 2023. It wasn't released yet, I think this one might take a while.

[0] https://ai.meta.com/blog/generative-ai-text-images-cm3leon/


I don't think they will ever release it considering it's likely much worse than flux.


> Any chance of this being released open weights?

Considering that Facebook/Meta releases blog posts titled "Open Source AI Is the Path Forward" but then refuses to actually release any Open Source AI, I'm guessing the answer is a hard "No".

They might release it under usage restrictions though, like they did with Llama, although probably only the smaller versions, to limit the output quality.


They have released a ton of open source? Llama 3 includes open training code, datasets, and models. Not to mention open-sourcing the foundation of most AI research today, pytorch.


Llama 3 is licensed under "Llama 3 Community License Agreement" which includes restrictions on usage, clearly not "Open Source" as we traditionally know it.

Just because pytorch is Open Source doesn't mean everything Meta AI releases is Open Source, not sure how that would make sense.

Datasets for Llama 3 is "A new mix of publicly available online data.", not exactly open or even very descriptive. That could be anything.

And no, the training code for Llama 3 isn't available, response from a Meta employee was: "However, at the moment-we haven't open sourced the pre-training scripts".


Sure, the Llama 3 Community License agreement isn't one of the standard open licenses and sucks that you can't use it for free if you're an entity the size of Google.

Here is the Llama source code, you can start training more epochs with it today if you like: https://github.com/meta-llama/llama3/blob/main/llama/model.p...

It's rumored Llama 3 used FineWeb, but you're right that they at least haven't been transparent about that: https://huggingface.co/datasets/HuggingFaceFW/fineweb

For models I prefer the term "open weight", but to assert they haven't open sourced models at all is plainly incorrect.


> Here is the Llama source code

Correct me if I'm wrong, but that's the code for doing inference?

Meta employee told me just the other day: "However, at the moment-we haven't open sourced the pre-training scripts", can't imagine they would be wrong about it?

https://github.com/meta-llama/llama-recipes/issues/693

> For models I prefer the term "open weight"

Personally, "Open" implies I can download them without signing an agreement with LLama, and I can do whatever I want with it. But I understand the community seems to think otherwise, especially considering the messaging Meta has around Llama, and how little the community is pushing back on it.

So Meta doesn't allow downloading the Llama weights without accepting the terms from them, doesn't allow unrestricted usage of those weights, doesn't share the training scripts nor the training data for creating the model.

The only thing that could be considered "open" would be that I can download the weights after signing the terms. Personally I wouldn't make the case that that's "open" as much as "possible to download", but again, I understand others understand it differently.


The source I linked is the PyTorch model, should be all you need to run some epochs. IDK what the pretraining scripts are.


Doesn't the training script need to have a training loop at least? Loss calculation? A optimizer? The script you linked contains neither, pretty sure that's for inference only


Oof you're right - no loss function or optimizer in place, so you'd need add that plus pull in data + tokenizer to get a training loop going.

Apologies - you are right and I was wrong. I would edit my comments but they're past the edit window, will leave a comment accordingly.


Past the edit window - want it to be higher up that only the model architecture is shared, no training scripts, as diggan correctly points out.


That and the NFSW finetunes that will inevitably follow; unlike the text-gen finetunes these could really cause trouble with deepfakes.


Deepfakes are already a reality, the technology is already there and good enough for harm, the genie is not going back to the bottle.

In fact, the more realistic the deepfakes become, the less harmful actual revenge porn and stolen sex videos can be, because of plausible deniability.


We live in a world where you can just say dumb bullshit about Haitians and millions of people will insist it's real.

This "good deepfakes will prevent harm because of plausible deniability" is absurd copium, and utterly divorced from reality.

Speak to victims some time. You are not helping them.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: