Black Forest Labs, in collaboration with mimic, has released FLUX-mimic, a new video-action model built on the FLUX 3 backbone. This release integrates advanced multimodal capabilities into robotics, focusing on video and action prediction.
FLUX 3 expands upon its predecessors by supporting multimodal content generation, specifically audio-visual content. It serves as the foundation for FLUX-mimic, which is designed to control robots through video-action modeling. The model has been tested and deployed in industrial settings such as Audi's production lines, demonstrating its applicability in real-world automation tasks.
Technically, FLUX 3 is trained jointly across images, video, and audio, with video prediction accounting for the majority of its compute costs. This extensive training allows the model to understand complex physical interactions, which is crucial for realistic video generation and robot control. The model integrates actions into its existing framework without a permanent loss of performance, indicating a shared backbone for video and action tasks.
FLUX-mimic is targeted at industries that require advanced robot learning and deployment capabilities. Potential users include developers and enterprises involved in robotics, particularly those focusing on dexterous manipulation and industrial automation. The model's ability to understand and predict complex physical interactions makes it suitable for a range of automation tasks.
Work implications: FLUX-mimic could augment roles in industrial automation and robotics development, enabling more efficient production processes and advanced robotic capabilities.
Originally reported by Black Forest Labs