An industry open letter titled “Open-Weights and American AI Leadership,” with 35 signatories including Nvidia, Microsoft, Meta, IBM, and Hugging Face, has triggered a very interesting debate on the use and necessity of open-weight AI models.
The letter states, “Open-weight models—AI models that anyone can download, inspect, modify, and run on their own infrastructure—are an important part of that foundation because they make advanced AI more accessible, adaptable, and widely available.” Basically, this U.S. led movement advocates open-weight AI models to expand access, foster competition, and enhance security by enabling broad community testing, adaptation, and ownership. It argues that while risks exist, transparency improves safety, and policy should support compute access, shared assets, and balanced regulation. Central to this is that open models drive innovation, sovereignty, and prosperity, avoiding restrictions that hinder progress.

The situation is complex, especially in our VFX space where there is a real cybersecurity issue of running bleeding-edge open-source GenAI workflows in-house. ComfyUI, while amazing, and a tool we have written about in the past, fxguide is quick to note that it can present a massive cybersecurity liability, as bad actors can slip viruses or worms into the application code or in through the open plugin ecosystem. There is just too much code to audit for an in-house IT team, so unless you are a big enough company to hire a team of software devs to forensically vet the GenAI solutions you could use. The alternative could be to build a trusted in-house solution, made with reputable open source ‘building block’ libraries. But, VFX industry security experts would normally treat general Open Source GenAI solutions as ‘hostile’, and that can particularly be a major issue with artists sometimes using their own devices, using their own downloaded tools.
Additionally, with the enormous shortage of GPUs and RAM prices having skyrocketed, some mid-level companies will still want to use cloud services. For those teams, while consumer GPUs and professional workstation GPUs can run Gen AI models, when scaling past one generation at a time, serious-sized teams are often quickly looking at data centres for the issue of scale.
But in the generative media space, open-weight models can offer other great opportunities beyond those mentioned in the open letter, which can greatly facilitate integrating generative AI (GenAI) tools into media and entertainment workflows. As we have started working with OpenWeight models in the fxguide tech bunker, we were keen to speak with GenAI workflow and colour science consultant JD Vandenberg (ex-Disney and Netflix) who wrote the following piece discussing the issue.

JD Vandenberg:
Adopting open-weight models facilitates workflow stability, security, cost predictability, and delivery efficiency. Although it requires technical expertise and a bring-your-own-infrastructure (BYOI) posture, it also enables deeper workflow integration by leveraging existing infrastructure instead of paying for external cloud services
Workflow stability
During production, workflow stability and predictability are paramount. No software gets updated unless necessary. In the GenAI space, models are updated often without warnings or full disclosure changelogs. An image, video, or audio track inferred with the exact same input and seed may change unless the provider offers versioned snapshots and you pin one.
Worse, a GenAI tool can disappear overnight, jeopardizing an entire production. That’s what happened to the Critterz project when OpenAI shut down the Sora app and web service in March. Critterz missed its planned Cannes debut, and its co-producers had to find a new AI partner, now targeting a release in early 2027. The same announcement ended Disney’s announced $1 billion investment in OpenAI and the three-year character licensing agreement attached to it; a deal that had been public for three months and was never finalized.
Security
There is still a lot of scepticism and fear that anything you feed the model could be used for training purposes. Production networks are segmented, with egress tightly controlled under MPA and TPN content security requirements. Sending source material to a hosted model and pulling generated assets back opens a path through the very controls content security exists to enforce. Every one of those paths widens the attack surface, and every one has to be justified to someone.
The material moving through that path is usually pre-release assets. Legal agreements help, but it is a remedy after the fact, not a control. At some point, someone has to establish that the provider’s infrastructure can actually hold that material, which means either paying for an audit or trusting a vendor that has never been through one. Studios already have a framework for this, yet only a few GenAI providers have completed a TPN assessment. On-prem models can run even in fully air-gapped networks without transferring any data to a third party.
Cost predictability
AI generation is highly iterative, and it is still very hard to predict how quickly a model will “get it right.”
A good example comes from Air Head, the 2024 Sora short made by the Toronto collective Shy Kids. In an interview with fxguide, post supervisor Patrick Cederberg estimated that the roughly ninety seconds of finished footage came out of “hundreds of generations at 10 to 20 seconds a piece,” and put the ratio at around 300:1 between source material and final cut. He flagged his own arithmetic as loose, and the two figures don’t reconcile precisely: take the conservative reading, and it is still north of 80:1. Scripted narrative typically shoots somewhere between 10:1 and 30:1, which is often accurately estimated in advance from experience.
As of August 2026, API rates of roughly $0.05 to $0.75 per second of output mean the cost of a finished minute is not a per-second price; it is a per-second price multiplied by a ratio difficult to forecast. The software industry has already reckoned with this. In June, a single employee at a San Francisco fintech ran up $81,267 in tokens in one week, while building a game: an illustration of how iterative, agentic workflows can turn token pricing into a material operational risk. Generative media runs the same loop at a considerably higher unit price. Running models on-prem removes the anxiety of the prompt slot machine: the possibility of waking up to a six-figure bill for a sequence that still isn’t approved. The hardware costs what it costs, whether you generate ten takes or ten thousand.
The obvious objection is to set a spend cap. Many platforms offer budgets or spend controls, but such a cap only relocates the problem. Imagine that the meter stops mid-sequence, on a delivery date, and you are negotiating a budget increase at 2 a.m. instead of finishing the shot. On-prem doesn’t make generation cheap, but it makes the cost bounded and known before the show starts. That’s a cost structure closer to what a production schedule actually requires.
Format and delivery
The gap between what GenAI tools produce and what a post pipeline ingests is not just a file size problem. It is primarily a format mismatch.
Most generative video tools deliver H.264: 8-bit, 4:2:0, display-referred (often sRGB). A low bit depth and dynamic range, compressed format without alpha channels, depth information, cryptomattes, scene-linear data, and featuring banding that survives every downstream grade. Cederberg noted that Sora offered no way to render additional passes such as mattes or depth at all, and that Shy Kids generated everything at 480p and upscaled externally. A VFX delivery is frequently EXR: a single 4K half-float frame at three channels is roughly 53MB uncompressed, and multi-layer EXRs of 200MB per frame are not unusual (that’s 576GB for a 2min shot at 24fps). To be clear, this issue is being addressed, but there is still a large gap to close.
At those rates, cloud storage and egress become material costs and transfer time becomes a scheduling constraint. But the deeper issue is that a model running in someone else’s cloud emits whatever format is cheapest for them to serve, while a model running on your own hardware can sit inside your own color pipeline and workflow and write what your pipeline actually reads.
Leveraging existing infrastructure
Most video generation models ship distilled or quantized variants that run on a single workstation GPU. The trend is running in our favor: quantization and distillation have been pushing VRAM floors down faster than parameter counts push them up. I can run quantized and distilled variants of most current open-weight video models on a single RTX 5090 (a consumer GPU with 32GB of VRAM). At the top end, an RTX PRO 6000 Blackwell with 96GB of VRAM retails around $13,250, though it launched materially cheaper.
Worth remembering, too, that studios, VFX houses, and post facilities already own substantial GPU capacity. The render farms and fleets of high-end GPU workstations are not idle infrastructure waiting to be bought; they are infrastructure waiting to be pointed at a different workload. For instance, a color corrector (e.g., FilmLight’s Baselight or Blackmagic DaVinci Resolve, often running on multiple GPUs) is ideally built for media inference instead of sitting idle outside of normal working hours.
Workflow Integration
The last advantage is the hardest to demonstrate in a demo reel and the most important in practice: a model you host is a model you can integrate with the tools your team already uses. That means invoking generation from a Nuke node or a Resolve plugin rather than a browser tab. It means dispatching jobs to the existing render manager, with the same queue, the same priorities, and the same logs as every other render on the show. It means color management that is declared once in an OCIO config or ACES pipeline rather than guessed at per clip, and version tracking that lives in the same asset management system as everything else. A hosted API can be scripted against, but it sits outside your pipeline looking in. A model you host lives inside it.
Open-Weights do not mean Open Source
One caution before you build a pipeline around any of this. “Open-weights” is not a single thing. It means you can download the parameters, but it does not automatically mean you can fine-tune, redistribute, or ship the output in a commercial delivery. Open source, however, usually means the software’s source code is available under an OSI-style license that allows people to inspect, modify, and redistribute it.
The frontier open-weight video and audio models sit at meaningfully different points on that spectrum. Some are Apache 2.0. Some carry community licenses with use restrictions. At least one requires a separate commercial agreement. A few are non-commercial only. A model that produces beautiful results under a license that forbids commercial use is a model you cannot put in a delivery, and the conform is an expensive place to discover that.
To summarize, models can be open-weight, but not open source (downloadable model files with a restrictive community license), and vice versa. Most leading open-weight models in this space come from outside the US, and policy around open-weight models is currently in flux. It is a fast-moving space and worth watching, and it’s another argument for archiving the weights you actually use rather than pulling them from a hub at job time.
Conclusion
Closed-weight models accessible through a web interface or API aren’t going away. Besides the benefits listed above, running models on-prem is far from being trivial. Many creators and organizations do not (or should not) burden themselves with running their own infrastructure. However, open-weight models offer new ways to fit better the existing, time-proven workflows that media and entertainment companies have spent decades optimizing. They also enable smoother scaling and more predictable costs. Models are getting bigger, while frontier model creators are finding creative ways to fit them into consumer-grade GPUs. Open weights don’t mean they’re free to use commercially, and Linux is a great example. You can run Fedora or CentOS Stream free of charge, but support comes with a Red Hat Enterprise Linux (RHEL) subscription. Some frontier model companies (LTX and Black Forest Labs, for instance) have adopted a similar business model, and I wouldn’t be surprised if more follow.

