
Background removal performance analysis — Citrix ctxbeffect & Fluendo flubkgndremoval

Written by
Aleix FigueresSeptember 8, 2026
Table of Contents
- 1. Overview
- 2. Benchmark Scope
- 3. Processing Architecture and GStreamer Pipelines
- 4. Measurement Methodology
- 5. Performance Results
- 6. Runtime Stability
- 7. Visual Quality Analysis
- 8. Conclusions
1. Overview
This report compares two background-removal solutions, Citrix ctxbeffect and Fluendo FAIP v1.2.6 flubkgndremoval, focusing on processing performance, latency, resource consumption, and output quality under the same benchmark framework.
FAIP v1.2.6 integrates Raven AI Engine v0.4.7 and provides GPU-native processing capabilities when configured with CUDA inference and Vulkan rendering. In the evaluated GPU-native path, frames are provided through CUDAMemory, processed with Vulkan, inferred with CUDA, and exported as VulkanImage. By keeping compatible stages GPU-addressable, this architecture is designed to reduce unnecessary host/device transitions and take advantage of GPU parallelism.
The objective of this benchmark is to quantify the practical performance characteristics of both solutions and evaluate the processing characteristics of FAIP’s GPU-native architecture under real-time and high-load conditions.
The benchmark evaluates how these architectural differences translate into throughput, processing latency, resource usage, scalability, and visual quality. From a deployment perspective, higher processing capacity and lower latency provide additional headroom for more demanding real-time video-processing workloads and can be relevant for thin-client and high-density video-processing scenarios.
1.1 Executive Summary
This campaign focuses on three representative execution paths covering the main CPU and GPU deployment options evaluated in the benchmark.
Representative execution paths:
Citrix
└── CPU/raw -> CPU processing -> CPU effect -> CPU postprocessing -> CPU/raw
FAIP CPU
└── CPU/raw -> CPU processing -> CPU inference -> CPU postprocessing -> CPU/raw
FAIP GPU-native
└── CUDAMemory -> Vulkan processing -> CUDA inference -> Vulkan postprocessing -> VulkanImage
These three paths provide the main product-level comparison used throughout the Executive Summary: Citrix running on CPU, FAIP running on CPU, and FAIP using its GPU-native CUDA+Vulkan path. In FAIP, processing and postprocessing/render stages can use CUDA or Vulkan depending on the selected GPU-native execution path. The remaining FAIP execution permutations are analyzed in Section 3.
Scenario A and Scenario B answer two complementary questions:
- Scenario A: input is paced to the target 30 or 60 FPS rate, representing real-time operation and reported as Observed FPS and Delivery Ratio.
- Scenario B: pacing is removed to measure maximum processing capacity, reported as Throughput FPS and Headroom.
Scenario A — Real-Time Performance
Scenario A evaluates each execution path with input paced to the target 30 or 60 FPS rate, showing observed behavior under real-time operating conditions.
Real-Time Output
| Workload | Citrix | FAIP CPU | FAIP GPU-native |
|---|---|---|---|
| 720p (30/60 FPS) | 30.00 / 24.00 | 30.00 / 60.00 | 27.50 / 52.00 |
| 1080p (30/60 FPS) | 17.00 / 22.50 | 30.00 / 60.00 | 28.00 / 34.50 |
| 1440p (30/60 FPS) | 28.00 / 7.50 | 30.00 / 28.00 | 16.00 / 28.00 |
| 4K (30/60 FPS) | 5.50 / 1.50 | 17.50 / 18.00 | 15.00 / 29.00 |
Under paced real-time conditions, FAIP provides stronger output at several demanding workloads. However, the GPU-native path does not reach the target input rate in all tested configurations, with the observed delivery ratio decreasing for some higher-resolution and higher-framerate workloads. This indicates that not all expected frames were observed at the final output during the measurement interval, but does not by itself demonstrate that frames were dropped by the FAIP plugin.
Scenario B — Maximum Processing Performance
Scenario B removes real-time pacing to measure the maximum processing capacity of each execution path and the corresponding plugin latency and resource utilization under load.
Processing Throughput
| Workload | Citrix | FAIP CPU | FAIP GPU-native |
|---|---|---|---|
| 720p (30/60 FPS) | 30.00 / 60.50 | 30.50 / 60.00 | 159.50 / 141.00 |
| 1080p (30/60 FPS) | 30.00 / 59.50 | 30.00 / 39.50 | 135.50 / 112.50 |
| 1440p (30/60 FPS) | 30.00 / 40.50 | 28.00 / 27.50 | 95.00 / 112.00 |
| 4K (30/60 FPS) | 16.00 / 19.00 | 19.50 / 16.00 | 85.00 / 161.50 |
In unpaced processing, the GPU-native FAIP path provides the largest headroom across the evaluated workloads, showing the practical benefit of the CUDA+Vulkan execution path when real-time pacing is removed.
Direct Plugin Latency
| Workload | Citrix | FAIP CPU | FAIP GPU-native |
|---|---|---|---|
| 720p (30/60 FPS) | 7.62 / 11.57 | 20.51 / 10.92 | 4.90 / 5.38 |
| 1080p (30/60 FPS) | 10.73 / 10.30 | 11.15 / 18.59 | 5.43 / 6.40 |
| 1440p (30/60 FPS) | 13.79 / 13.56 | 21.98 / 23.20 | 6.87 / 5.90 |
| 4K (30/60 FPS) | 33.84 / 29.02 | 23.92 / 30.21 | 7.01 / 3.57 |
Resource Usage
| Workload | Citrix CPU | Citrix RAM | FAIP CPU CPU | FAIP CPU RAM | FAIP GPU-native CPU | FAIP GPU-native RAM |
|---|---|---|---|---|---|---|
| 720p (30/60 FPS) | 30.9% / 103.0% | 81.4 MB / 80.5 MB | 1129.8% / 1394.4% | 541.3 MB / 543.9 MB | 898.5% / 808.5% | 986.8 MB / 953.6 MB |
| 1080p (30/60 FPS) | 49.1% / 93.3% | 105.6 MB / 106.3 MB | 1194.2% / 1271.6% | 554.9 MB / 556.0 MB | 793.0% / 725.4% | 939.9 MB / 926.7 MB |
| 1440p (30/60 FPS) | 73.0% / 98.9% | 140.6 MB / 140.1 MB | 1204.2% / 1208.6% | 578.5 MB / 578.4 MB | 712.7% / 688.6% | 951.2 MB / 934.1 MB |
| 4K (30/60 FPS) | 102.2% / 117.2% | 213.5 MB / 224.8 MB | 973.2% / 803.3% | 609.0 MB / 614.3 MB | 680.4% / 1017.6% | 969.4 MB / 1074.8 MB |
FAIP GPU-native combines lower direct plugin latency with substantially higher processing headroom, while Citrix maintains a lighter host CPU/RAM footprint across the same workloads.
Visual Quality Analysis
| Sample A - Citrix | Sample B - Fluendo |
|---|---|
![]() | ![]() |
Both solutions provide effective foreground/background separation, but with different visual styles: Citrix emphasizes smoother blur, while FAIP uses opaque/replaced background rendering in this configuration. At a high level, both solutions show similar subject-preservation behavior, with differences mainly in background treatment aesthetics.
1.2 Key Findings
Maximum processing capacity: FAIP GPU-native delivers the highest Scenario B throughput, reaching 159.50 FPS at 720p30 and 161.50 FPS at 4K60.
Lower plugin latency: FAIP GPU-native provides the lowest direct plugin latency, with 6.40 ms at 1080p60 and 3.57 ms at 4K60, versus 10.30 ms and 29.02 ms for Citrix.
Real-time performance: FAIP performs strongly under paced workloads. At 1080p60, FAIP CPU reaches 60 FPS versus 22.50 FPS for Citrix. At 4K60, FAIP GPU-native provides the highest observed output at 29 FPS.
Better high-load scaling: FAIP GPU-native maintains substantially more processing headroom as resolution and framerate increase, particularly in Scenario B.
Resource trade-off: Citrix has a significantly lighter CPU/RAM footprint, while FAIP uses more resources in exchange for higher processing capacity and lower GPU-native latency.
Runtime stability: FAIP still shows teardown instability in some execution paths, which remains an area for further hardening.
Overall, the benchmark shows two distinct deployment profiles. Citrix offers the lighter CPU-oriented footprint, while FAIP GPU-native provides the strongest processing-capacity and latency profile when GPU resources are available. FAIP CPU provides a useful CPU-backed execution option, but at significantly higher host-resource cost. The preferred architecture therefore depends on whether the deployment prioritizes minimum resource footprint or maximum processing headroom and latency performance.
2. Benchmark Scope
2.1 Benchmark Matrix
| Dimension | Value |
|---|---|
| Configurations | 9 |
| Resolutions | 4 (720p, 1080p, 1440p, 4K) |
| Target framerates | 2 (30, 60) |
| Scenarios | 2 |
| Total | 9 x 4 x 2 x 2 = 144 |
Workloads: 720p30, 720p60, 1080p30, 1080p60, 1440p30, 1440p60, 4K30, 4K60.
Case counts:
- Scenario A: 72
- Scenario B: 72
- FAIP: 128
- Citrix: 16
2.2 Benchmark Environment
| Component | Configuration |
|---|---|
| OS | Ubuntu 22.04 |
| GPU | NVIDIA GeForce RTX 3090 |
| GStreamer | 1.28.2 |
| FAIP | v1.2.6 |
| Raven AI Engine | v0.4.7.0 |
| Citrix plugin | ctxbeffect |
| FAIP plugin | flubkgndremoval |
| AI device classes | CPU / CUDA |
| Render path classes | CPU / CUDA / Vulkan |
2.3 Evaluated Configurations
| ID | Requested execution family |
|---|---|
faip_cpuraw_cpu_cpu | CPU/raw -> CPU processing -> CPU AI -> CPU/raw |
faip_cpuraw_cpu_cuda | CPU/raw -> CPU processing -> CUDA AI -> CPU/raw |
faip_cpuraw_vulkan_cpu | CPU/raw -> Vulkan processing -> CPU AI -> VulkanImage |
faip_cpuraw_vulkan_cuda | CPU/raw -> Vulkan processing -> CUDA AI -> VulkanImage |
faip_cudamemory_cuda_cpu | CUDAMemory -> CUDA processing -> CPU AI -> CPU/raw |
faip_cudamemory_cuda_cuda | CUDAMemory -> CUDA processing -> CUDA AI -> CPU/raw |
faip_cudamemory_vulkan_cpu | CUDAMemory -> Vulkan processing -> CPU AI -> VulkanImage |
faip_cudamemory_vulkan_cuda | CUDAMemory -> Vulkan processing -> CUDA AI -> VulkanImage |
citrix_cpu_cpu_cpu | CPU/raw -> CPU processing -> CPU AI -> CPU/raw |
Stale theoretical IDs such as faip_cudamemory_cpu_cpu and faip_cudamemory_cpu_cuda are excluded from this corrected final matrix.
2.4 Benchmark Scenarios
For a fixed workload/configuration, both scenarios preserve the same architecture and content family; pacing mode is the principal experimental difference.
2.4.1 Scenario A - Live Stream / Real-Time Performance
Scenario A uses paced input at the target nominal framerate (30 or 60 FPS) to show observed real-time output behavior.
+----------------------+ +----------------------+ +--------------------------+
| Input paced 30/60 FPS|-->| Selected exec path |-->| Observed: FPS + delivery |
+----------------------+ +----------------------+ +--------------------------+
2.4.2 Scenario B - Offline / Maximum Processing Performance
Scenario B removes input pacing to measure maximum processing capacity and associated latency/resource behavior under load.
+----------------------+ +----------------------+ +----------------------------------+
| Input unpaced |-->| Selected exec path |-->| Measured: thrpt+latency+host res |
+----------------------+ +----------------------+ +----------------------------------+
3. Processing Architecture and GStreamer Pipelines
The GStreamer pipeline describes inter-element topology and memory exchange. The FAIP Updated execution plan describes the internal processing/render and AI backend actually selected at runtime.
3.1 Processing Architecture
High-level pipeline architecture:
Input
|
v
Decoder
|
v
Memory Carrier
|
+--------------------------+
| |
v v
CPU/raw CUDAMemory
| |
v v
Background Removal Plugin
|
v
Output Memory Carrier
|
v
Sink
Conceptual plugin internal architecture:
Background Removal Plugin
|
v
+------------------+
| Pre-processing |
| CPU/CUDA/Vulkan |
+------------------+
|
v
+------------------+
| AI Inference |
| CPU / CUDA |
+------------------+
|
v
+------------------+
| Post-processing |
| CPU/CUDA/Vulkan |
+------------------+
|
v
Output Frame
3.2 Execution-Path Summary Table
This subsection summarizes the evaluated execution configurations in a compact table.
| Configuration ID | Input family | Processing backend | AI backend | Output family | Upload/Download | Representative execution path |
|---|---|---|---|---|---|---|
citrix_cpu_cpu_cpu | CPU/raw (Y4M) | CPU | CPU effect | CPU/raw | no / no | CPU/raw -> CPU processing -> CPU effect -> CPU/raw |
faip_cpuraw_cpu_cpu | CPU/raw (Y4M) | CPU | CPU | CPU/raw | no / no | CPU/raw -> CPU processing -> CPU AI -> CPU/raw |
faip_cpuraw_cpu_cuda | CPU/raw (Y4M) | CPU | CUDA | CPU/raw | no / no | CPU/raw -> CPU processing -> CUDA AI -> CPU/raw |
faip_cpuraw_vulkan_cpu | CPU/raw (Y4M) | Vulkan | CPU | VulkanImage | yes / no | CPU/raw -> Vulkan processing -> CPU AI -> VulkanImage |
faip_cpuraw_vulkan_cuda | CPU/raw (Y4M) | Vulkan | CUDA | VulkanImage | yes / no | CPU/raw -> Vulkan processing -> CUDA AI -> VulkanImage |
faip_cudamemory_cuda_cpu | CUDAMemory (H.264 + NVDEC) | CUDA | CPU | CPU/raw | no / yes | CUDAMemory -> CUDA processing -> CPU AI -> CPU/raw |
faip_cudamemory_cuda_cuda | CUDAMemory (H.264 + NVDEC) | CUDA | CUDA | CPU/raw | no / yes | CUDAMemory -> CUDA processing -> CUDA AI -> CPU/raw |
faip_cudamemory_vulkan_cpu | CUDAMemory (H.264 + NVDEC) | Vulkan | CPU | VulkanImage | yes / no | CUDAMemory -> Vulkan processing -> CPU AI -> VulkanImage |
faip_cudamemory_vulkan_cuda | CUDAMemory (H.264 + NVDEC) | Vulkan | CUDA | VulkanImage | yes / no | CUDAMemory -> Vulkan processing -> CUDA AI -> VulkanImage |
Scenario A and Scenario B use the same execution-path selection for a given configuration; the primary experimental difference is input pacing mode.
4. Measurement Methodology
4.1 Throughput and Plugin Latency
Throughput uses the final sink rendered-frame rate over the measurement window. Scenario A reports observed FPS with delivery context; Scenario B reports throughput with headroom against nominal workload FPS.
plugin sink timestamp
|
v
background-removal plugin
|
v
plugin src timestamp
PluginLatency_i = t_src,i - t_sink,i
Direct plugin latency is measured from matched sink/src PTS using monotonic timestamps. Latency is stage-local and is not equivalent to 1000/FPS or pure CUDA kernel time.
4.2 Resource Measurement and Execution Validation
CPU mean [%], RAM mean [MB], GPU utilization [%], and GPU memory [MB] are sampled in-window. CPU can exceed 100% because process CPU is summed across logical cores.
Execution-path interpretation uses this evidence hierarchy: FAIP Updated execution plan, negotiated caps, AI-device selection, and Vulkan environment. measurement_validity and runtime_stability remain independent dimensions, so VALID + TEARDOWN_ABORT can occur when teardown fails after measurement completion.
5. Performance Results
Primary comparison set in this section:
- Citrix CPU reference:
citrix_cpu_cpu_cpu - FAIP CPU reference:
faip_cpuraw_cpu_cpu - FAIP GPU-native reference:
faip_cudamemory_vulkan_cuda
Data-integrity audit policy:
- All cells were verified against testcase-level rows in
results_smoke_enriched.jsonfor exact scenario, workload, configuration, validity, FPS, latency, and resources. - If any primary row were invalid, the table cell would be
N/E; no focused rerun values are substituted into the primary matrix.
5.1 Real-Time Performance
Scenario A paced real-time results (Observed FPS with delivery ratio):
| Workload | Citrix CPU | FAIP CPU | FAIP GPU-native |
|---|---|---|---|
| 720p30 | 30.00 FPS (100.0%) | 30.00 FPS (100.0%) | 27.50 FPS (91.7%) |
| 720p60 | 24.00 FPS (40.0%) | 60.00 FPS (100.0%) | 52.00 FPS (86.7%) |
| 1080p30 | 17.00 FPS (56.7%) | 30.00 FPS (100.0%) | 28.00 FPS (93.3%) |
| 1080p60 | 22.50 FPS (37.5%) | 60.00 FPS (100.0%) | 34.50 FPS (57.5%) |
| 1440p30 | 28.00 FPS (93.3%) | 30.00 FPS (100.0%) | 16.00 FPS (53.3%) |
| 1440p60 | 7.50 FPS (12.5%) | 28.00 FPS (46.7%) | 28.00 FPS (46.7%) |
| 4K30 | 5.50 FPS (18.3%) | 17.50 FPS (58.3%) | 15.00 FPS (50.0%) |
| 4K60 | 1.50 FPS (2.5%) | 18.00 FPS (30.0%) | 29.00 FPS (48.3%) |
5.2 Maximum Processing Performance
Scenario B unpaced maximum processing results (Throughput FPS with headroom):
| Workload | Citrix CPU | FAIP CPU | FAIP GPU-native |
|---|---|---|---|
| 720p30 | 30.00 FPS (1.00x) | 30.50 FPS (1.02x) | 159.50 FPS (5.32x) |
| 720p60 | 60.50 FPS (1.01x) | 60.00 FPS (1.00x) | 141.00 FPS (2.35x) |
| 1080p30 | 30.00 FPS (1.00x) | 30.00 FPS (1.00x) | 135.50 FPS (4.52x) |
| 1080p60 | 59.50 FPS (0.99x) | 39.50 FPS (0.66x) | 112.50 FPS (1.88x) |
| 1440p30 | 30.00 FPS (1.00x) | 28.00 FPS (0.93x) | 95.00 FPS (3.17x) |
| 1440p60 | 40.50 FPS (0.68x) | 27.50 FPS (0.46x) | 112.00 FPS (1.87x) |
| 4K30 | 16.00 FPS (0.53x) | 19.50 FPS (0.65x) | 85.00 FPS (2.83x) |
| 4K60 | 19.00 FPS (0.32x) | 16.00 FPS (0.27x) | 161.50 FPS (2.69x) |
5.3 Resolution and Framerate Scalability
This subsection interprets scaling trends from Sections 5.1 and 5.2.
In Scenario A (paced), Observed FPS and Delivery Ratio generally decline as resolution and target framerate increase. CPU-reference paths and GPU-native behavior are both workload-dependent, and GPU-native FAIP does not consistently reach the Scenario A target input rate at higher-demand points. Below-target Scenario A output indicates fewer frames observed at final output, but it does not by itself demonstrate plugin-side frame drops.
In Scenario B (unpaced), Throughput FPS and Headroom separate the execution architectures more clearly: CPU-reference paths become constrained earlier at demanding resolution/framerate combinations, while FAIP GPU-native retains substantially more processing headroom across the workload range.
Refer to Section 5.1 for Scenario A detailed values and Section 5.2 for Scenario B detailed values.
5.4 Plugin Latency
Scenario A mean/p95 direct plugin latency (ms), all 8 workloads:
| Workload | Citrix Mean / p95 | FAIP CPU Mean / p95 | FAIP GPU-native Mean / p95 |
|---|---|---|---|
| 720p30 | 13.753 / 14.455 | 10.466 / 13.525 | 4.479 / 5.968 |
| 720p60 | 13.560 / 14.228 | 10.453 / 12.162 | 3.944 / 5.504 |
| 1080p30 | 18.112 / 19.352 | 10.709 / 13.268 | 4.783 / 8.103 |
| 1080p60 | 10.768 / 14.864 | 10.879 / 12.684 | 6.605 / 12.247 |
| 1440p30 | 13.455 / 15.930 | 12.902 / 15.357 | 7.817 / 12.929 |
| 1440p60 | 14.136 / 17.293 | 23.139 / 35.442 | 7.462 / 14.249 |
| 4K30 | 31.974 / 36.283 | 26.234 / 38.151 | 9.544 / 16.705 |
| 4K60 | 29.967 / 33.575 | 25.980 / 41.124 | 7.807 / 13.683 |
Scenario B mean/p95 direct plugin latency (ms), all 8 workloads:
| Workload | Citrix Mean / p95 | FAIP CPU Mean / p95 | FAIP GPU-native Mean / p95 |
|---|---|---|---|
| 720p30 | 7.623 / 10.200 | 20.507 / 37.332 | 4.896 / 9.142 |
| 720p60 | 11.573 / 14.050 | 10.919 / 11.550 | 5.385 / 10.140 |
| 1080p30 | 10.735 / 13.064 | 11.155 / 12.515 | 5.430 / 9.923 |
| 1080p60 | 10.303 / 12.890 | 18.587 / 28.296 | 6.403 / 12.533 |
| 1440p30 | 13.795 / 16.820 | 21.981 / 29.431 | 6.869 / 12.723 |
| 1440p60 | 13.561 / 16.445 | 23.197 / 31.872 | 5.901 / 11.858 |
| 4K30 | 33.839 / 36.692 | 23.924 / 34.410 | 7.007 / 15.213 |
| 4K60 | 29.020 / 37.858 | 30.208 / 41.595 | 3.571 / 5.057 |
5.5 Resource Utilization
CPU utilization (%):
| Workload | Citrix A | FAIP CPU A | FAIP GPU-native A | Citrix B | FAIP CPU B | FAIP GPU-native B |
|---|---|---|---|---|---|---|
| 720p30 | 59.5 | 1170.9 | 486.2 | 30.9 | 1129.8 | 898.5 |
| 720p60 | 67.8 | 1385.0 | 638.1 | 103.0 | 1394.4 | 808.5 |
| 1080p30 | 71.0 | 1183.8 | 522.7 | 49.1 | 1194.2 | 793.0 |
| 1080p60 | 59.6 | 1411.9 | 227.8 | 93.3 | 1271.6 | 725.4 |
| 1440p30 | 69.6 | 1190.8 | 131.6 | 73.0 | 1204.2 | 712.7 |
| 1440p60 | 74.8 | 1205.7 | 140.8 | 98.9 | 1208.6 | 688.6 |
| 4K30 | 93.9 | 823.1 | 134.0 | 102.2 | 973.2 | 680.4 |
| 4K60 | 130.0 | 874.0 | 125.2 | 117.2 | 803.3 | 1017.6 |
RAM (MB):
| Workload | Citrix A | FAIP CPU A | FAIP GPU-native A | Citrix B | FAIP CPU B | FAIP GPU-native B |
|---|---|---|---|---|---|---|
| 720p30 | 80.6 | 577.4 | 1053.1 | 81.4 | 541.3 | 986.8 |
| 720p60 | 79.5 | 541.8 | 1051.3 | 80.5 | 543.9 | 953.6 |
| 1080p30 | 104.7 | 556.0 | 1079.7 | 105.6 | 554.9 | 939.9 |
| 1080p60 | 106.3 | 555.5 | 918.9 | 106.3 | 556.0 | 926.7 |
| 1440p30 | 140.0 | 579.5 | 929.5 | 140.6 | 578.5 | 951.2 |
| 1440p60 | 139.3 | 578.3 | 926.7 | 140.1 | 578.4 | 934.1 |
| 4K30 | 223.4 | 615.2 | 966.1 | 213.5 | 609.0 | 969.4 |
| 4K60 | 210.4 | 613.5 | 982.7 | 224.8 | 614.3 | 1074.8 |
GPU utilization (%):
| Workload | Citrix A | FAIP CPU A | FAIP GPU-native A | Citrix B | FAIP CPU B | FAIP GPU-native B |
|---|---|---|---|---|---|---|
| 720p30 | 3.0 | 2.0 | 3.0 | 0.5 | 4.5 | 15.5 |
| 720p60 | 4.0 | 4.5 | 6.0 | 1.5 | 5.5 | 11.0 |
| 1080p30 | 3.0 | 3.0 | 4.0 | 2.5 | 7.5 | 14.5 |
| 1080p60 | 1.5 | 5.5 | 5.0 | 0.0 | 6.0 | 13.5 |
| 1440p30 | 0.0 | 9.5 | 4.0 | 4.0 | 12.0 | 11.5 |
| 1440p60 | 2.0 | 11.5 | 3.5 | 3.0 | 9.5 | 13.5 |
| 4K30 | 3.0 | 7.5 | 4.5 | 3.5 | 4.0 | 11.5 |
| 4K60 | 4.0 | 7.0 | 3.0 | 3.5 | 5.0 | 17.5 |
GPU memory (MB):
| Workload | Citrix A | FAIP CPU A | FAIP GPU-native A | Citrix B | FAIP CPU B | FAIP GPU-native B |
|---|---|---|---|---|---|---|
| 720p30 | 271.0 | 622.0 | 992.0 | 535.0 | 966.0 | 1253.0 |
| 720p60 | 271.0 | 618.0 | 992.0 | 420.0 | 755.0 | 1253.0 |
| 1080p30 | 271.0 | 686.0 | 1046.0 | 535.0 | 1138.5 | 1317.0 |
| 1080p60 | 535.0 | 682.0 | 1301.0 | 535.0 | 971.0 | 1315.0 |
| 1440p30 | 535.0 | 1042.5 | 1391.0 | 535.0 | 977.0 | 1379.0 |
| 1440p60 | 535.0 | 1095.0 | 1411.0 | 538.0 | 1095.0 | 1411.0 |
| 4K30 | 535.0 | 1287.0 | 1645.5 | 535.0 | 1111.0 | 1651.5 |
| 4K60 | 414.0 | 1287.0 | 1647.5 | 271.0 | 1287.0 | 1386.0 |
6. Runtime Stability
6.1 Runtime Stability Summary
| Runtime status / metric | Count | Interpretation |
|---|---|---|
| Total executed | 144 | Complete benchmark matrix |
| Initial successful measurements | 142/144 | Valid results in original campaign |
| Initial zero-frame anomalies | 2 | Original measurement anomalies |
| STABLE | 51 | Measurement and teardown completed normally |
| TEARDOWN_ABORT | 91 | Measurement completed, but process aborted during teardown |
| Configuration mismatches | 0 | Requested execution path matched |
| Negotiation failures | 0 | No caps-negotiation failure |
7. Visual Quality Analysis
7.1 Static Analysis (Reference Frames)
7.1.1 Citrix Reference Frame

7.1.2 FAIP Reference Frame

In static-frame comparison, both solutions show clear human segmentation/matting quality. Citrix appears more strict at subject borders, while FAIP appears slightly smoother in edge treatment. Despite these style differences, final perceived quality is similar.
7.2 Dynamic Analysis (Processed Output Animations)
7.2.1 Citrix Processed Output

7.2.2 FAIP Processed Output

In dynamic clips, the same conclusion holds: both pipelines maintain effective subject separation over time, with Citrix preserving stricter border definition and FAIP keeping a smoother visual transition. Motion behavior appears stable for both in these examples, and overall output quality remains comparable.
8. Conclusions
The benchmark shows that FAIP’s main performance advantage comes from its GPU-native CUDA+Vulkan architecture. When input pacing is removed in Scenario B, faip_cudamemory_vulkan_cuda consistently provides the largest processing headroom and lowest direct plugin latency of the representative execution paths, demonstrating the benefit of keeping video processing, AI inference, and frame exchange on GPU-addressable memory paths.
Under real-time paced conditions (Scenario A), the results are more workload-dependent. FAIP CPU provides strong delivery at several 720p and 1080p workloads, while the GPU-native path does not consistently reach the target input rate despite its much higher maximum processing capacity in Scenario B. At the most demanding 4K60 workload, however, FAIP GPU-native provides the highest observed output of the three representative paths. This difference between Scenario A and Scenario B also shows that maximum processing capacity alone does not guarantee equivalent paced real-time delivery, and that pacing, synchronization, memory flow, and pipeline behavior must be considered together.
The performance advantage of FAIP comes with a clear resource trade-off. Citrix maintains a substantially lighter CPU and RAM footprint, whereas FAIP—particularly its GPU-backed configurations—uses more host and GPU resources. The two solutions therefore represent different deployment profiles: Citrix favors a lightweight CPU-oriented execution model, while FAIP GPU-native favors processing headroom and low plugin latency when GPU resources are available.
From a scalability perspective, the results support the architectural direction of FAIP: as workload complexity increases, the CUDA+Vulkan GPU-native path retains a substantially wider processing envelope than the CPU-reference paths in maximum-load operation. This makes the GPU-native architecture particularly relevant where processing capacity and latency are more important than minimizing resource consumption.
Finally, runtime stability remains the main technical area requiring further hardening. Teardown-abort and memory-corruption signatures were observed even when measurements themselves completed successfully. These issues do not invalidate completed performance measurements, but they should be resolved before treating the evaluated paths as fully mature for long-running production workloads.
Overall, FAIP GPU-native provides the strongest performance profile in this benchmark, particularly for maximum processing capacity and direct plugin latency, while Citrix provides the more resource-efficient baseline. The results also identify two clear priorities for FAIP: improve paced real-time delivery consistency and resolve the observed teardown instability, while preserving the performance advantages of the GPU-native execution path.
8.1 Technical Q&A
Why can CPU usage exceed 100%?
Process CPU usage is aggregated across logical cores, so multi-threaded pipelines can report values above 100%.
Why can Scenario B FPS exceed Scenario A for the same workload?
Scenario A is paced by design, while Scenario B removes pacing to expose maximum processing capacity.
Does CUDA inference alone mean the full pipeline is GPU-native?
No. Input carrier, processing backend, inference backend, and output carrier are distinct dimensions. The runtime execution plan is authoritative.
How should VALID + TEARDOWN_ABORT be interpreted?
It indicates measurement completed successfully, but the process later aborted on teardown.
Are these results long-duration sustainability claims?
No. They are comparative benchmark measurements and are not long-duration endurance evidence.
Ready to put it to the test? Download and try Fluendo AI Plugins 🚀
For further information, get in touch with our team to discuss how we can help optimize your custom GStreamer and GPU pipelines.
