BFL releases FLUX 3 Action: a 7B robot model
Summary
Black Forest Labs has released FLUX 3 Action, a collection of 7B AI models for robotics, including base models and policies for SO-101 and DROID, available on Hugging Face.
View Cached Full Text
Cached at: 09/23/26, 07:01 PM
FLUX 3 Action - a black-forest-labs Collection
Source: https://huggingface.co/collections/black-forest-labs/flux-3-action
![]()
black-forest-labs’s Collections
updatedabout 2 hours ago
FLUX 3 Action base and shared encoders, SO-101 policy, and DROID policy with optimization variants.
-
—
#### black-forest-labs/flux-3-action-base Robotics• Updatedabout 6 hours ago • 21 NoteAction adaptation base and shared frozen video/text encoders. -
—
#### black-forest-labs/flux-3-action-so101 Robotics• 7B• Updatedabout 6 hours ago • 48 • 10 NoteSO-101 LeRobot policy with saved processors and task-LoRA recipe. -
—
#### black-forest-labs/flux-3-action-droid Robotics• 7B• Updatedabout 3 hours ago • 33 • 9 NoteValidated DROID policy at root; opt-in optimization variants in subfolders.
Similar Articles
BFL Introduces FLUX 3 - multi-modal model for Image, Video and Audio
BFL has introduced FLUX 3, a multi-modal AI model capable of generating images, videos, and audio.
Flux 3 X Mimic: The Next Generation of Video-Action Models
Black Forest Labs announces FLUX 3, a multimodal foundation model that jointly generates audio-visual content and, via collaboration with mimic robotics, enables video-action prediction for robot control, tested at Audi.
Black Forest Lab's Flux 3: Omni-modality for image, video, audio & action prediction
Black Forest Lab's Flux 3 is a new omni-modal AI model capable of generating and predicting images, video, audio, and actions.
black-forest-labs/FLUX.1-dev
Black Forest Labs releases FLUX.1-dev, a 12-billion parameter open-weights text-to-image transformer model, available on Hugging Face with API endpoints and local inference support.
FLUX 3
FLUX 3 is a multimodal AI model from Black Forest Labs that generates synchronized video and audio from text prompts with optional image or video inputs.