IBC 2026 recap: C2PA, WebRTC, and the next era of broadcasting
broadcastingIBC 2026 was a huge success! Explore our recap on Fluendo's latest live video tech, our Content Everywhere presentation, and the future of multimedia.
Keep your products updated with Fluendo’s broadcast solutions to cover your engineering requirements. Bring the latest audio and video codecs and avoid the hassle of compliance with a specific standard.


our case studies
Developments that bring real-world results, these case studies show how our solutions help your business achieve goals and enhance user experiences.

The client needed to replace manual and error-prone ad spotting with a high-accuracy AI-powered advertisement detection system that could run privately at the edge. Key requirements included precise ad recognition in both video and audio, support for multiple languages, and offline processing to minimize bandwidth usage, protect sensitive multimedia content, and reduce operational costs.
The solution was delivered as a self-contained Docker application with a full CLI, ensuring portability, easy maintenance, and scalability. Designed for robustness and future-readiness, this AI-driven advertisement detection solution empowers businesses to automate ad tracking while maintaining efficiency, privacy, and cost control.
The client transitioned from manual ad spotting to an AI-powered advertisement detection system achieving 96.36% average accuracy with minimal timing offsets across English, Spanish, Basque, Catalan, and Galician. By deploying AI at the edge, the solution eliminates constant cloud dependency, reducing costs, accelerating processing, and ensuring full control of sensitive multimedia content. This not only simplifies compliance but also minimizes manual effort and enables confident verification.
In practice, the system streamlines multimedia operations, strengthens advertiser relationships, and scales seamlessly with demand. Built to support reliable, multilingual ad detection technology, the solution reflects Fluendo’s commitment to delivering future-ready AI solutions for multimedia content analysis.
Achieved 96.36% accuracy in ad recognition
Eliminated cloud dependency by deploying AI at the edge
Scaled seamlessly with demand via Docker portability
Automated ad detection eliminates manual reviewing to drastically speed up media tracking
Achieved 96.36% accuracy in ad recognition
Eliminated cloud dependency by deploying AI at the edge
Scaled seamlessly with demand via Docker portability
Automated ad detection eliminates manual reviewing to drastically speed up media tracking

The company is a technology integrator specialized in complex engineering projects, delivering scalable video infrastructure for government and defense systems. As part of a major public-sector initiative, Fluendo expanded the client’s Recasting & Recording System (RRS) to support real-time video stream ingestion, advanced recasting and recording formats, video transcoding, and asynchronous event notifications. Leveraging a modular GStreamer-based architecture, these enhancements were seamlessly integrated in a short timeframe, ensuring high performance, long-term scalability, and reliability for mission-critical multimedia operations.
Fluendo efficiently delivered the requested features, positioning the RRS as a potential replacement for the client’s existing solution. Robust testing ensured high reliability and performance.
Supported real-time ingestion for major public-sector initiatives
Ensured high reliability for mission-critical defense systems
Seamlessly added transcoding and asynchronous notifications
Delivered a high-performance replacement for legacy solutions
Supported real-time ingestion for major public-sector initiatives
Ensured high reliability for mission-critical defense systems
Seamlessly added transcoding and asynchronous notifications
Delivered a high-performance replacement for legacy solutions

The company develops a modern AI-powered camera system designed to create safer workplaces and enable smarter business operations across industries. To strengthen its multimedia engineering capabilities, the company partnered with Fluendo to enhance its team’s expertise in GStreamer with a strong focus on Rust-based development.
Recognizing Rust’s growing adoption for its memory safety and high performance, Fluendo delivered a six-day, hands-on training program combining GStreamer fundamentals with modern Rust practices. Participants were guided through pipeline creation, memory-safe plugin development, interoperability with existing C APIs, and debugging complex multimedia workflows within the Rust ecosystem. The progressive structure of the sessions enabled engineers to directly apply their learnings to the company’s AI video processing products.
The engineering team gained confidence in integrating Rust into their multimedia stack with GStreamer. They developed a strong foundation in safe, high-performance multimedia development, enabling them to create more robust and efficient multimedia systems internally.
Leveraged Rust’s memory safety for AI video processing
Deep dive into C API interoperability and the Rust ecosystem
Deep dive into C API interoperability and the Rust ecosystem
Rust-based GStreamer pipelines eliminate memory bugs to guarantee crash-free performance
Leveraged Rust’s memory safety for AI video processing
Deep dive into C API interoperability and the Rust ecosystem
Deep dive into C API interoperability and the Rust ecosystem
Rust-based GStreamer pipelines eliminate memory bugs to guarantee crash-free performance

The Catalan audiovisual sector needed professional-grade synthetic voices, but existing open-source models lacked the naturalness and prosody required for high-end production.
Furthermore, AI voice cloning posed significant legal and ethical risks, including intellectual property infringement and unauthorized use of biometric data.
A secure, end-to-end multimedia solution was required to close this technological gap while ensuring compliance with the EU AI Act and protecting voice talent. The objective was to build a highly scalable, ethical platform that provided industry-leading phonetic precision without compromising data sovereignty.

To address the project’s complex requirements, the solution was divided into two distinct technological pillars: a highly optimized TTS model and a secure, enterprise-grade web application.
To ensure professional naturalness, the system used Transfer Learning from existing foundational models (Proyecto AINA), which were heavily refined using over 45 hours of studio-mastered recordings. A critical differentiator was the manual linguistic correction of the phonetizer engine, which drastically reduced pronunciation errors and ensured perfect dialectal accuracy. The architecture deployed a modular TTS engine that combined eSpeak-NG, Matcha-TTS, and Vocos. This rigorous acoustic and linguistic optimization successfully reduced the Phoneme Error Rate (PER) to just 0.95%. The application and model were presented to a focus group of the audiovisual industry, and the experts rated it as the best Catalan TTS available.
Surrounding the AI engine, a robust application was built to manage the complexities of voice actors’ IP management through contracts, GDPR compliance, and high-concurrency requests. Orchestrated via a backend using Django, Redis, and Celery, the entire pipeline guarantees absolute data sovereignty. Crucially, the platform addressed the legal challenges of Generative AI by embedding C2PA cryptographic metadata into every generated audio file, ensuring the synthetic media’s origins were immutable and protected against unauthorized data mining.
Supported by real-time ASR validation on the frontend, this scalable architecture achieved an internal synthesis latency of just 0.12 seconds, a Real-Time Factor (RTF) of 0.17, and a 100% success rate for backend inference requests. Ultimately, the project delivered a transparent, ethical, and high-performance solution ready for the professional dubbing industry.
Achieved 0.95% Phoneme Error Rate (PER) and high MOS
Ensured compliance with the EU AI Act and GDPR
Protected voice talent through C2PA traceability
Adapted to specific phonetic nuances for regional dubbing
Achieved 0.95% Phoneme Error Rate (PER) and high MOS
Ensured compliance with the EU AI Act and GDPR
Protected voice talent through C2PA traceability
Adapted to specific phonetic nuances for regional dubbing

Fraunhofer HHI required the integration of Digitally Signed Content (DSC) support into GStreamer to ensure content integrity and traceability across multimedia workflows. At the time, GStreamer did not provide standardized mechanisms for embedding and validating digitally signed metadata within encoded video streams, which created a gap for organizations seeking trusted and verifiable media pipelines.
Fluendo’s expertise was essential because of its 20-year history as a leader in the GStreamer community and its previous success with the SPIRIT OC1 project, which served as the baseline for this engagement. As multimedia experts with deep insight into the framework, the team was uniquely positioned to implement DSC SEI messages on top of H.266/VVC bitstreams. This project aimed to deliver a modular, maintainable solution that could be contributed upstream to the open-source community.

The diagnostic process identified that traditional authentication methods relied on container-level signatures, which were easily stripped or manipulated. To address this, the solution focused on embedding cryptographic signatures directly into the video bitstream at the Network Abstraction Layer (NAL) unit level. This ensured that authentication information traveled with the video data itself, making it significantly harder to tamper with.
The implemented solution involved extending the H.266 multimedia parsers to extract SEI messages carrying DSC metadata and expose them as GstMeta. New GStreamer elements were developed to handle DSC signing and verification, allowing for the generation of signing metadata in compliance with JVET standards and the validation of content using public keys. These elements emitted signals to notify the system of validation success or failure, providing a robust framework for real-time media authentication.
The engagement resulted in a modular implementation that was contributed upstream to the GStreamer community, ensuring long-term maintainability and industry adoption. Further technical details were shared in a related blog post.
Aligned with JVET-AK0194 and ITU H.274 standards
Enabled content authenticity and anti-tampering layers
Developed modular elements for signing and verifying metadata
Embedded cryptographic signatures directly into the video bitstream
Aligned with JVET-AK0194 and ITU H.274 standards
Enabled content authenticity and anti-tampering layers
Developed modular elements for signing and verifying metadata
Embedded cryptographic signatures directly into the video bitstream
our use cases
These use cases present conceptual examples of how our ideas and technologies could address real-world industry challenges.

As video becomes central to entertainment, social media, e-learning, and news reporting, creators increasingly film in uncontrolled or sensitive environments—city streets, classrooms, hospitals, protests, and corporate settings where bystanders may appear without consent. Under GDPR and similar data-protection laws, publishing identifiable individuals without permission risks legal action and reputational harm. An automated anonymization system can detect and obscure faces in live streams or recordings without undermining creative vision.
Recent advancements in AI and real-time processing now allow scalable anonymization pipelines for post-production, livestreaming, and mobile content capture.
Anonymizer is a Fluendo AI Plugin built to fit seamlessly into today’s media production workflows, whether editing on a desktop, capturing on mobile, or livestreaming from the field. Powered by a GStreamer pipeline, it delivers real-time face blurring and frame-accurate anonymization across various video formats, codecs, and resolutions.
Designed for content creators, journalists, educators, broadcasters, and corporate teams, Anonymizer enables privacy-safe video creation in both live and post-production scenarios. Its scalable architecture supports on-device and centralized processing, allowing you to automatically anonymize video content anywhere—with minimal setup and no manual intervention—including dedicated features for children anonymization to meet stricter privacy regulations.
Protect identities without cloud processing. Film in sensitive areas without post-clearance delays — ensuring privacy, compliance, and faster production.
Delivers GDPR-aligned workflows by anonymizing personal data before it is stored or analyzed, ensuring regulatory compliance.
Shields against takedowns, lawsuits, and platform penalties.
Protect identities without cloud processing. Film in sensitive areas without post-clearance delays — ensuring privacy, compliance, and faster production.
Delivers GDPR-aligned workflows by anonymizing personal data before it is stored or analyzed, ensuring regulatory compliance.
Shields against takedowns, lawsuits, and platform penalties.

Sports clubs and academies increasingly produce live video streams, interviews, commentary shows, and behind-the-scenes content for digital platforms and social media. These broadcasts often take place in training grounds, stadiums, or mixed media areas, where children, staff members, and spectators may appear in the background.
This creates important privacy and safeguarding challenges, particularly in youth sports environments where minors must not be publicly identifiable without explicit consent.
Manual editing or post-production anonymization is not feasible for live broadcasts or real-time streaming, where video must be processed instantly before distribution.
With advances in AI-based person and face detection, video processing systems can automatically identify individuals appearing in the background of a live stream and apply anonymization techniques such as face blurring or masking in real time. This allows sports organizations to safely broadcast interviews, live shows, and training content while protecting the identity of children and other individuals present in the scene.
Our solution uses deep learning models to detect individuals appearing in live video streams, including players, staff members, spectators, and children present in the background during interviews, commentary segments, or live coverage.
Each video frame is analyzed in real time to identify visible faces or people in the scene. Once detected, the system automatically applies privacy-preserving transformations such as face blurring, pixelation, or masking, ensuring that individuals cannot be identified while preserving the visual integrity of the broadcast.
The system is designed to operate within high-performance live production environments, supporting video resolutions up to 4K and 8K, and high frame rates including 60 fps and 120 fps. Thanks to an ultra-optimized multimedia AI pipeline, anonymization can be performed with very low latency, ensuring that the broadcast workflow remains uninterrupted.
The solution integrates directly into professional streaming and broadcast pipelines, enabling anonymization to occur before encoding or distribution to streaming platforms. Processing can be deployed on edge AI devices located in stadiums or production units, or in cloud-based infrastructures, depending on the production workflow.
By combining optimized AI inference with high-performance video processing, sports organizations can safely produce interviews, analyst shows, live streams, and behind-the-scenes content while protecting the identities of children, spectators, and staff—even in high-resolution, high-frame-rate live productions running on compact edge hardware.
Complies with GDPR during live global broadcasts
Prevents accidental privacy violations in real-time
Delivers seamless anonymization without affecting the viewer experience
Implements complex AI processing at live-broadcast speeds
Complies with GDPR during live global broadcasts
Prevents accidental privacy violations in real-time
Delivers seamless anonymization without affecting the viewer experience
Implements complex AI processing at live-broadcast speeds

The Catalan audiovisual sector currently faces a significant technological gap regarding high-quality, professional-grade synthetic voice solutions. Foundation Text-to-Speech (TTS) technologies work effectively for the majority of languages with large numbers of speakers, such as English, Chinese, and Spanish. Nevertheless, for regional languages like Catalan, these models do not deliver adequate performance and fail to meet the needs of the phonetic nuances and dialectal diversity required for professional dubbing and media production. Furthermore, the rise of generative AI has raised urgent concerns about the privacy of biometric data and the intellectual property rights of voice actors.
Due to these challenges, the industry would require a secure, sovereign, and ethically grounded platform to generate high-fidelity Catalan synthetic speech. It would be essential to establish a system that not only achieves naturalness through advanced Computer Vision and MLOps principles but also ensures total compliance with the EU AI Act and GDPR.
Catalan serves as a proof of concept for sovereign technology in other underrepresented European regional languages.
The proposed solution would consist of a scalable MLOps pipeline to train Sovereign Foundation Models for regional underserved EU languages. To establish the Catalan prototype, the highly specialized voice cloning stage would utilize Low-Rank Adapters (LoRA) to separate the foundational style from specific vocal identities.
This parameter-efficient approach would allow the rapid injection of unique Catalan timbres without retraining the entire core engine. Ultimately, this Catalan-first approach would guarantee high-end audiovisual production while maintaining complete data sovereignty.
Ensures voice production is handled ethically and transparently
Delivers high-fidelity synthetic voices for professional media
Applies AI to regional languages with high phonetic accuracy
Conforms to professional audiovisual production standards
Ensures voice production is handled ethically and transparently
Delivers high-fidelity synthetic voices for professional media
Applies AI to regional languages with high phonetic accuracy
Conforms to professional audiovisual production standards

As the digital broadcast landscape transitions toward the IP-based ATSC 3.0 standard (NextGen TV), the industry faces a critical infrastructure gap. Unlike legacy systems, this new framework utilizes the same Internet Protocol backbone as major OTT platforms, enabling a hybrid model of over-the-air signals and interactive internet content.
Within this ecosystem, Dolby AC-4 has emerged as the mandatory audio standard, offering advanced capabilities such as immersive Atmos metadata, dialog enhancement, and personalized audio streams.
The transition to Next Generation Audio (NGA) is moving from a premium feature to a global regulatory requirement: United States (ATSC 3.0 / NextGen TV), European Union (DVB-T2 & UHD), Brazil (TV 3.0), South Korea (UHD services).
The rapid adoption of AC-4 by major networks for live sports and high-profile events has created significant technical hurdles. Current hardware and monitoring tools risk obsolescence without a high-fidelity, legal decoding path. Furthermore, legacy decoders are unable to process the complex metadata required for modern audience measurement and immersive audio experiences. The reliance on “experimental” or uncertified decoders in commercial cloud-encoding SaaS environments introduces substantial legal and technical risks for broadcasters and hardware OEMs.

The proposed solution would bridge the “legal gap” by providing the first certified decoding path for the industry’s most critical multimedia frameworks: GStreamer and FFmpeg. By integrating a certified fluac4dec GStreamer plugin and FFmpeg Enabler+Dolby Proffesional, software-defined receivers, gateways, and monitoring tools would gain the ability to accurately “unzip” highly compressed AC-4 data and transform it into actionable high-fidelity audio.
Complies with the latest broadcasting standards
Provides superior audio decoding and immersive sound
Meets the rigorous technical requirements of the broadcasting industry
Integrates next-gen audio technology for future-proof solutions
Complies with the latest broadcasting standards
Provides superior audio decoding and immersive sound
Meets the rigorous technical requirements of the broadcasting industry
Integrates next-gen audio technology for future-proof solutions
Bits & Bytes
Explore our blog, one byte at a time. Our team unpack our latest news, industry insights and in-depth articles to connect you with the multimedia world.Blog
Read more about our work
IBC 2026 recap: C2PA, WebRTC, and the next era of broadcasting
broadcastingIBC 2026 was a huge success! Explore our recap on Fluendo's latest live video tech, our Content Everywhere presentation, and the future of multimedia.
Raven AI Engine v0.4.7: From an AI inference engine to a heterogeneous GPU computing platform
broadcasting, video-surveillance, multimedia-edge-ai, gstreamer, fluendo-ai-plugins, anonymizer, ravenRaven v0.4.7 and FAIP v1.2.6 extend GPU-accelerated AI processing to Ubuntu x64 with CUDA, Vulkan, heterogeneous execution, and user-defined shaders.
Real-time face avatarization: identity swapping on the edge
broadcasting, video-surveillance, multimedia-edge-ai, application-development, outsourceReal-time, AI-powered face avatarization that swaps a chosen identity onto live video, running entirely on the edge.
Face anonymization with real-time re-identification
broadcasting, video-surveillance, multimedia-edge-ai, application-development, outsourceAutomated, AI-powered face anonymization with real-time re-identification, running entirely on the edge.