Open Source and Open Weight:
a distinction you cannot afford to ignore

Open Source
and Open Weight:
a distinction you cannot afford to ignore

In the public debate on artificial intelligence, terms overlap, blur and — often not by accident — are used interchangeably, even by those who have every interest in doing so. One of the most common misconceptions concerns the difference between an Open Source model and an Open Weight model. A seemingly technical distinction, but one with concrete implications for anyone using these tools in a professional or commercial context.

What Open Source actually means

The concept of open source originates in the software world and has a precise definition, established by the Open Source Initiative (OSI): software is open source when its source code is freely accessible, modifiable and redistributable, without significant restrictions on commercial use and without discrimination against individuals, groups or fields of application.

In other words, open source does not simply mean “I can download it for free”.
It means I can study it, modify it, integrate it into a commercial product, distribute it in a modified version — all without having to ask permission or pay royalties.

This definition, consolidated over decades, is the one many people have in mind when they hear about “open” AI models.

What Open Weight models actually are

With the spread of large language models, a different category has emerged — one often presented to the public under the open source label, but substantially different: **open weight** models.

In this case, what is made public are the model weights — the mathematical parameters that define the system’s behaviour after training. Downloading the weights means being able to run the model locally, without depending on the producer’s infrastructure.

So far, it sounds like Open Source. The problem lies in the licence.

Open Weight models are distributed under licences that impose often significant restrictions:

  • prohibition on using the model to train competing systems
  • limitations on commercial use beyond certain thresholds
    – obligation to attribute the original model name in derivative products
    – clauses prohibiting specific fields of application
  • obligation to attribute the original model name in derivative products
  • clauses prohibiting specific fields of application

The complete training source code, moreover, is rarely made available. The weights are there, but the recipe for recreating the model from scratch — the data, the procedures, the architectural choices — usually remains proprietary.

The most prominent case: Meta and the Llama family

L’esempio più discusso è quello di Llama, la famiglia di modelli sviluppata da Meta®. Llama è stata presentata pubblicamente come una scelta a favore dell’apertura e della ricerca, e i suoi pesi sono scaricabili liberamente. Questo ha portato molti a considerarla una soluzione open source.

The reality is more nuanced. Early versions of Llama included explicit restrictions on commercial use for organisations above a certain size, and a prohibition on using the model to train other AI systems. Subsequent versions have relaxed some of these restrictions, but the licence remains distant from the classical definition of open source.

The OSI itself has taken a position on the matter, excluding Llama and similar models from its own definition of open source, precisely because of the lack of full transparency on training data and usage limitations.

The other side of the problem:
what happens to what you upload

The other side of the problem: what happens to what you upload

Up to this point we have discussed rights over outputs — what the AI produces. But for a company using these tools in a professional context, the mirror question is at least equally critical: what happens to its own data when used as input?

Consider a concrete case

A fashion brand working on a collection not yet presented to the market might think of using sketches, sample photographs or prototype images as references to generate visualisations, mood boards or graphic variations through an AI tool.

From a Creative standpoint, it makes sense. From a Confidentiality standpoint, the risk is significant. .

Every cloud AI platform manages user data according to its own terms of service. Some explicitly reserve the right to use inputs — texts, images, files — to train or improve their models. Others allow users to opt out of this practice through specific settings. Others still, in their enterprise versions, guarantee complete data isolation. But how many professionals read these terms carefully before uploading confidential material?

For a fashion brand, an unreleased collection is a competitive asset with a temporary and precisely defined window of value. The competitive advantage exists until the moment of official presentation: before that moment, any uncontrolled dissemination — even involuntary, even mediated by a third-party platform — can nullify it. The same applies to advertising campaigns not yet launched, product concepts, positioning strategies, internal documents.

Once data has been uploaded to a cloud system not directly controlled by the company, control over that data is effectively transferred — within the limits and modalities established by a licence that is rarely read with the attention it deserves.

This introduces a further practical distinction between cloud models and locally executable models. An open weight model installed on one’s own infrastructure — with all the licensing complexities we have discussed — guarantees by definition that inputs remain within the corporate perimeter. It is one of the concrete reasons why some organisations choose on-premise solutions despite the greater technical complexity and management costs.

The choice between a cloud service and a local solution is therefore not merely a question of performance or cost: it is also a question of intellectual property and data confidentiality.

Why this distinction matters for those working in the industry

For a communications agency, a creative studio or a professional integrating AI tools into their workflow, the difference is not academic.

Using an Open Weight model in a commercial project for a client without verifying the licence terms creates concrete exposure. The same logic applies to outputs: the technical ability to generate content does not automatically coincide with the right to use it commercially.

The parallel with the stock photography world is immediate. A royalty-free image is not an image without rights — it is an image with rights granted more broadly than under a rights-managed model, but still within limits defined by the licence. In the same way, an “open” model is not a model without constraints.

The strategy behind the ambiguity

There is a reason why this distinction tends to remain blurred in the public communications of many AI companies.

Presenting a model as “open” generates trust, widespread adoption and a developer community that builds around it. Over time, that community — and the products it has created — becomes an asset difficult for competitors to replicate. The model producer, meanwhile, retains control over the most advanced versions, the cloud infrastructure and enterprise agreements.

It is a legitimate business model, but it is worth understanding it for what it is: a distribution strategy with precise commercial objectives, not an ideological choice in favour of knowledge sharing.

There is no single official logo to recognize an Open Weight license because Open Weight is a descriptive classification for AI models, not a single standard or specific legal entity.

What to verify before integrating an AI model into a professional Workflow

Regardless of how a model is presented, it is worth verifying several points before adopting it in a commercial context:

Does the licence explicitly permit commercial use?
**Does the licence explicitly permit commercial use?**
The fact that a model is freely downloadable is not sufficient.

Can the outputs be used without restrictions?
Some licences limit not only the use of the model but also the rights over generated outputs.

Are there size or sector limitations?
Certain models exclude categories of users — companies above a revenue threshold, specific sectors, particular fields of application.

Can the model be integrated into a product sold to third parties?
This is often the most critical clause for those working on commission.

Is the data uploaded as input used to train the model?
Images, documents, sketches, unpublished texts: before uploading confidential material to any AI platform, it is essential to verify how inputs are handled. For sensitive or unreleased content, consider solutions with guaranteed data isolation or locally executable models.

The confusion between open source and open weight is not set to resolve itself any time soon, not least because none of the parties benefiting from the ambiguity has a direct interest in clarifying it. For those using these tools professionally, however, navigating licences with precision is an integral part of the work — exactly as it has been, and continues to be, in the world of images, music and any other content subject to rights.