Skip to content

Latest commit

 

History

History
2521 lines (1843 loc) · 192 KB

File metadata and controls

2521 lines (1843 loc) · 192 KB

Mangekyo — Multi-Graphics API GPU Benchmark Report


Table of Contents

Part 1 — How It Works

  1. Project Overview
  2. Pipeline Architecture
  3. RenderDoc Frame-Capture Analysis
  4. Python Benchmark Tooling

Part 2 — Primary GPUs: RX 9070 XT & RX 6900 XT

  1. RX 9070 XT and RX 6900 XT — Cross-API, Particle Scaling, and Headless
  2. Swapchain Throttling, Timestamp Pollution, and Headless Compute

Part 3 — AMD GPU Generational Analysis

  1. AMD GPU Generational Analysis — TeraScale 2 → GCN → RDNA 2 → RDNA 4
  2. Cross-Validation Against 3DMark

Part 4 — Other GPUs and Cross-Platform

  1. Cross-API Comparison — RTX 5090
  2. Cross-GPU Comparison — Vulkan
  3. Memory Allocation Impact — Vulkan, RTX 5090
  4. Software Renderer Baseline — WARP
  5. Legacy Discrete GPU vs Modern iGPU

Part 5 — Platform and API Issues

  1. OpenGL GPU Selection — Platform Limitations
  2. DX11 Timestamp Query Failures
  3. OpenGL Compute Shader Performance on AMD GPUs
  4. Dual Identical GPU Behaviour (Mac Pro 2013)

Part 6 — Conclusion

  1. Summary

Appendices


Part 1 — How It Works

1. Project Overview

What This Project Does

A cross-platform GPU compute and rendering microbenchmark written in C++17. It simulates millions of particles on the GPU using a compute shader, then renders them as point-sprites using a graphics pipeline — all within a single command buffer per frame.

The project implements five interchangeable graphics API backends:

Backend API Shader Language Platforms
Vulkan 1.1+ Explicit, low-level GLSL → SPIR-V Windows, Linux, macOS (MoltenVK)
DirectX 12 Explicit, low-level HLSL (SM 5.1) Windows 10+
DirectX 11 Implicit, driver-managed HLSL (SM 5.0) Windows 7+
OpenGL 4.3 Implicit, driver-managed GLSL 430 Windows, Linux
Metal Explicit (Apple) MSL macOS

All five backends share the same AppBase class for windowing (GLFW), particle initialisation, timing, and benchmark result management. Each backend overrides InitBackend(), DrawFrame(), CleanupBackend(), and WaitIdle().

Benchmark Workload Suite

Beyond the original particle simulation, the benchmark runs five interchangeable workloads, each isolating a different axis of GPU performance and reporting a deterministic, cross-API-comparable metric (not just FPS). Select with --workload; every backend runs the same algorithm so numbers compare directly.

Axis Workload (--workload) What it stresses Metric Design
Bandwidth stream (default) Memory subsystem GB/s Particle update — ~0.15 FLOP/byte, bandwidth-bound
Compute (achievable) nbody FP32 ALU + SFU + shared memory GFLOP/s All-pairs N-body, tiled through shared memory (--bodies)
Fill / fragment stress Rasteriser + fragment ALU/SFU + ROP G-iter/s Fullscreen fractal, fixed per-pixel iterations (--iter)
Compute (peak) synthpeak Raw ALU throughput per precision GFLOPS / GIOPS Register-resident FMA loop, vkpeak-style (--precision, --iter)
3D render render3d Vertex transform + rasterisation + fill + depth MQuad/s Perspective + orbiting camera + depth test; instanced camera-facing billboard quads (--particles)

Why both nbody and synthpeak? synthpeak measures the theoretical ceiling (near-peak FLOPS), while nbody measures achievable compute under real data dependencies and special-function (rsqrt) load. Together they bracket compute performance; stream and stress add the bandwidth and fill axes; render3d adds a real 3D graphics pipeline (the others render nothing or only a 2D pass).

Cross-API results (RTX 5090, reference run)

Bandwidth / Compute / Fill — same workload across APIs:

Workload Vulkan DX12 DX11 OpenGL
N-body (GFLOP/s, 64K bodies) 46,681 36,074 49,094¹ 53,798
Fractal (G-iter/s, 8000 iter) 3,681 3,450 3,445¹ 3,521
Render3D (MQuad/s, 1M billboards) 1,861 1,978 1,655 1,774

¹ DX11 figures from windowed mode (see DX11 timestamp note below).

The Render3D pass (instanced billboards + depth + perspective) converges across APIs much like the fill/peak axes — all four are within ~20%, since the work is dominated by hardware vertex/raster/fill throughput rather than API overhead.

Synthetic peak by precision (GFLOPS / GIOPS):

Precision Vulkan DX12 OpenGL DX11 Notes
FP32 112,615 109,016 98,687 109,737 ≈ 5090's ~105 TFLOPS spec
FP16 122,412 119,757 118,934 N/A² ≈ FP32×1.05 — consumer Blackwell has no 2× non-tensor FP16
FP64 1,980 1,973 1,971 1,667 ≈ 1/57 of FP32 (consumer 1/64 double rate)
INT32 61,508 56,978 61,557 63,147 ~0.55× FP32 (IMAD ≈ half-rate)

² DX11 cannot run true FP16: Direct3D 11 caps at Shader Model 5, whose min16float is minimum-precision (the driver may run it at FP32), not IEEE FP16. True 16-bit needs SM 6.2, which only the DX12 runtime exposes.

Key cross-API insight: the peak (synthpeak) and fill (stress) axes converge tightly across APIs (within ~6%) because they are pure hardware-throughput limited — the driver/API barely matters. The achievable compute (nbody) axis shows a larger spread (~1.5×, OpenGL fastest, DX12 slowest on this NVIDIA driver) because dispatch/scheduling overhead is exposed. Running multiple axes is what separates "hardware ceiling" from "API/driver overhead."

Synthetic peak by precision N-body achievable compute Fractal fill rate 3D render throughput

Precision support is bounded by each API, not by effort

FP16 has no single portable path, so each API uses the right tool for true 16-bit:

  • Vulkan — standard path: VK_KHR_shader_float16_int8 + GL_EXT_shader_explicit_arithmetic_types_float16, packed f16vec2. ✅
  • DX12 — FXC's min16float is only minimum-precision (often FP32). True FP16 needs Shader Model 6.2, so the FP16 kernel is precompiled to signed DXIL with the Windows SDK DXC (-enable-16bit-types) at build time and loaded at runtime; the device is checked for Native16BitShaderOps. ✅
  • OpenGL — desktop GL core has no portable FP16; the FP16 kernel uses GL_NV_gpu_shader5 (float16_t/f16vec2), gated at runtime by the extension (NVIDIA). AMD would use GL_AMD_gpu_shader_half_float. ✅ (NVIDIA)
  • DX11 — impossible: Direct3D 11 caps at SM 5, which has no true FP16. N/A.
  • Metal — half/half2 is native; written but untested here (no macOS). ✅ (code)

Likewise FP64 is unavailable on Metal — Apple GPUs have no double-precision units.

DX11 SynthPeak: kernel runs, but headless timing is unavailable

synthpeak (and the headless compute path generally) forces headless mode (no window/swapchain). On the test driver, DX11 never resolves GPU timestamp queries in headless mode — a pre-existing, documented behaviour (see woa-dx11-timestamp-issue.md). The DX11 SynthPeak kernel does execute and produce FPS, but with no GPU-time it cannot report a GFLOPS score, so it shows as N/A. DX11 timestamps work normally in windowed mode — which is why DX11's N-body and Fractal figures above (windowed) are present while its headless SynthPeak is not. This is a DX11/driver limitation, not an untested path.

Test Hardware

All benchmarks in this report were collected on four physical systems using the GPUs and CPUs listed below.

AMD GPUs (Primary Test Fleet)

Nine AMD GPUs spanning four architecture generations and 15 years (2009–2024), including seven discrete GPUs and two integrated GPUs:

GPU Architecture Year CU / SP FP32 (TFLOPS) Memory
RX 9070 XT RDNA 4 2025 64 CU (4096 SP) 48.70 16 GB GDDR6, 640 GB/s
RX 6900 XT RDNA 2 2020 80 CU (5120 SP) 23.04 16 GB GDDR6, 512 GB/s
RX 6600 XT RDNA 2 2021 32 CU (2048 SP) 10.60 8 GB GDDR6, 256 GB/s
Vega Frontier Edition GCN 5 (Vega) 2017 64 CU (4096 SP) 13.11 16 GB HBM2, 483 GB/s
RX 580 GCN 4 (Polaris 20) 2017 36 CU (2304 SP) 6.17 8 GB GDDR5, 256 GB/s
FirePro D700 GCN 1 (Tahiti) 2013 32 CU (2048 SP) 3.48 6 GB GDDR5, 264 GB/s
HD 5770 TeraScale 2 (Juniper) 2009 10 SIMD (800 SP) 1.36 1 GB GDDR5, 77 GB/s
Ryzen 7 9800X3D iGPU RDNA 2 2024 2 CU (128 SP) 0.56 Shared DDR5, ~90 GB/s
Ryzen 5 7600 iGPU RDNA 2 2023 2 CU (128 SP) 0.56 Shared DDR5, ~90 GB/s

RDNA 4 uses dual-issue FP32; traditional single-issue calculation yields 24.3 TFLOPS. The two iGPUs are architecturally identical (RDNA 2, 2 CU, 2200 MHz) but sit in different CPU platforms (Zen 5 3D V-Cache vs Zen 4).

NVIDIA Discrete GPUs

GPU Architecture Year CUDA Cores FP32 (TFLOPS) Memory
RTX 5090 Blackwell 2025 21,760 104.8 32 GB GDDR7, 1792 GB/s
GTX 970 Maxwell 2.0 2014 1,664 3.92 4 GB GDDR5, 224 GB/s

Qualcomm (Windows on ARM)

GPU Architecture Year ALU FP32 (TFLOPS) Memory
Adreno 640 Adreno 6xx 2019 768 ~0.90 Shared LPDDR4X, ~34 GB/s

Tested on a Xiaomi Pad 5 (Snapdragon 860) running Windows on ARM. Supports DX11 FL 11_1, DX12 FL 12_1, and Vulkan 1.1. Microsoft WARP (software renderer) is also used as a baseline — see Section 12.

Test Platforms (CPUs)

Also, All CPUs were tested with Microsoft WARP (DX11/DX12 software rasteriser) to measure CPU-side compute and rendering performance.

CPU Architecture Year Cores / Threads Boost Clock TDP Platform
AMD Ryzen 7 9800X3D Zen 5 + 3D V-Cache 2024 8C / 16T 5.2 GHz 120 W AM5, DDR5-6000
AMD Ryzen 5 7600 Zen 4 2023 6C / 12T 5.1 GHz 65 W AM5, DDR5-6000
Intel Xeon E5-2697 v2 Ivy Bridge-EP 2013 12C / 24T 3.5 GHz 130 W LGA 2011, DDR3-1866
Qualcomm Snapdragon 860 Kryo 485 (A76/A55) 2021 1P+3P+4E 2.96 GHz 2~7 W LPDDR4X, WoA

Key Features

  • Multi-GPU support: enumerate and select from all available GPUs (discrete → integrated → software), with --gpu CLI override.
  • GPU timestamp profiling: per-frame compute / render / total GPU timing via VkQueryPool (Vulkan), ID3D12QueryHeap (DX12), ID3D11Query (DX11), glQueryCounter (OpenGL).
  • Benchmark mode: fixed frame count or timed run, warmup period, min/max/avg statistics, CPU-bound vs GPU-bound analysis.
  • Result persistence: auto-save to ~/.gpu_bench/results.json, compare, delete, CSV export.
  • RenderDoc integration: VK_EXT_debug_utils labels + In-Application API for programmatic frame capture (--capture <frame>).
  • Python tooling: chart generation, batch benchmark automation, markdown/HTML report export, 3DMark cross-validation.

2. Pipeline Architecture

Per-Frame Execution Model

Every frame executes the following sequence in a single command buffer (shown for the Vulkan backend; other backends follow the same logical structure):

┌──────────────────────────────────────────────────────┐
│                  Command Buffer                       │
│                                                       │
│  ┌─ Timestamp T0 (TOP_OF_PIPE) ───────────────────┐  │
│  │                                                 │  │
│  │  ┌─ Particle Compute ────────────────────────┐  │  │
│  │  │  Bind compute pipeline                    │  │  │
│  │  │  Bind descriptor set (SSBO)               │  │  │
│  │  │  Push constants (deltaTime, damping)      │  │  │
│  │  │  vkCmdDispatch(N/256, 1, 1)               │  │  │
│  │  └───────────────────────────────────────────┘  │  │
│  │                                                 │  │
│  ├─ Timestamp T1 (COMPUTE_SHADER) ────────────────┤  │
│  │                                                 │  │
│  │  ┌─ SSBO Barrier ───────────────────────────┐  │  │
│  │  │  srcStage: COMPUTE_SHADER_BIT            │  │  │
│  │  │  dstStage: VERTEX_INPUT_BIT              │  │  │
│  │  │  SHADER_WRITE → VERTEX_ATTRIBUTE_READ    │  │  │
│  │  └──────────────────────────────────────────┘  │  │
│  │                                                 │  │
│  ├─ Timestamp T2 (TOP_OF_PIPE) ───────────────────┤  │
│  │                                                 │  │
│  │  ┌─ Particle Render ────────────────────────┐  │  │
│  │  │  Begin render pass (clear to dark blue)  │  │  │
│  │  │  Bind graphics pipeline                  │  │  │
│  │  │  Bind vertex buffer (same SSBO)          │  │  │
│  │  │  vkCmdDraw(particleCount, 1, 0, 0)       │  │  │
│  │  │  End render pass                         │  │  │
│  │  └──────────────────────────────────────────┘  │  │
│  │                                                 │  │
│  └─ Timestamp T3 (COLOR_ATTACHMENT_OUTPUT) ────────┘  │
│                                                       │
└──────────────────────────────────────────────────────┘

Compute Shader

The compute shader (shaders/compute.comp) performs Euler integration on each particle:

layout(local_size_x = 256) in;

layout(set = 0, binding = 0, std430) buffer ParticleBuffer {
    Particle particles[];   // vec4 position + vec4 velocity = 32 bytes
};

layout(push_constant) uniform ComputeParams {
    float deltaTime;
    float bounds;
};

void main() {
    uint i = gl_GlobalInvocationID.x;
    particles[i].position.xyz += particles[i].velocity.xyz * deltaTime;
    if (particles[i].position.x > bounds)
        particles[i].position.x = -bounds;
}
  • Workgroup size: 256 threads — balances occupancy across AMD (wavefront 64) and NVIDIA (warp 32) architectures.
  • SSBO layout: std430 guarantees C++-compatible packing (no padding).
  • Push constants: avoid descriptor set updates every frame; only 8 bytes pushed per dispatch.

For 16M particles: 16,777,216 / 256 = 65,536 workgroups dispatched.

Graphics Pipeline

The render pipeline draws all particles as POINT_LIST using the same SSBO as a vertex buffer:

Stage Configuration
Vertex Input Binding 0, stride 32 bytes: position (R32G32_SFLOAT, offset 0), colour (R32G32B32A32_SFLOAT, offset 16)
Vertex Shader Maps particle speed to a blue → red colour gradient, sets gl_PointSize = 2.0
Rasteriser Point topology, no culling
Fragment Shader Passes interpolated colour to output
Colour Blend Additive blending (SRC_ALPHA + ONE) — overlapping particles create bright clusters
Render Pass Single subpass, one colour attachment (swapchain image, B8G8R8A8_SRGB)

Memory Architecture

Mode Vulkan Flags When Used
Device-local (default) DEVICE_LOCAL + staging buffer copy Discrete GPU — particle data lives in VRAM
Host-visible (--host-memory) HOST_VISIBLE | HOST_COHERENT Integrated GPU or debugging — CPU-mappable, no staging copy

On discrete GPUs, the staging buffer is created, filled with initial particle data, copied via vkCmdCopyBuffer, then destroyed. All subsequent compute and render operations access only the device-local buffer.

Synchronisation

The single VkBufferMemoryBarrier between compute and render ensures:

  • All compute shader writes to particle positions are visible before the vertex shader reads them.
  • No additional barriers are needed because the entire frame is recorded into one command buffer on one queue.
  • On integrated GPUs with unified memory, the barrier is essentially a no-op (no cache flush between separate memory domains).

Timestamp Query Pipeline

Each backend implements GPU timestamp queries using a ring buffer of kTimestampSlotCount (8) frame slots to avoid blocking on in-flight frames:

API Write Read Sync Clock Frequency
Vulkan vkCmdWriteTimestamp vkGetQueryPoolResults Fence wait from previous frame timestampPeriod from device properties
DX12 ID3D12GraphicsCommandList::EndQuery ResolveQueryData + readback buffer Fence signal/wait ID3D12CommandQueue::GetTimestampFrequency
DX11 ID3D11DeviceContext::End(query) GetData with retry loop Disjoint query (S_FALSE → retry) D3D11_QUERY_DATA_TIMESTAMP_DISJOINT.Frequency
OpenGL glQueryCounter(GL_TIMESTAMP) glGetQueryObjectui64v GL_QUERY_RESULT_AVAILABLE poll Fixed 1 ns resolution

Four timestamps per frame yield three intervals: compute time (T1−T0), render time (T3−T2), and total GPU time (T3−T0).


3. RenderDoc Frame-Capture Analysis

Industry-standard GPU profiling of the Vulkan particle compute + render pipeline using RenderDoc — the most widely used cross-vendor, cross-API GPU frame debugger.

What is RenderDoc?

RenderDoc is a free, open-source (MIT licence) GPU frame debugger created by Baldur Karlsson. It intercepts a single frame's worth of graphics API calls, allowing post-mortem inspection of every resource, pipeline state, and GPU operation. It is one of the tools referenced by AMD's JD under "Use of industry-standard profiling and debug tools".

Tool Vendor Platform
RenderDoc Open-source Vulkan, DX11, DX12, OpenGL — all GPUs
PIX Microsoft DX12 (Windows only)
Radeon GPU Profiler (RGP) AMD Vulkan, DX12 (AMD GPUs only)
Nsight Graphics NVIDIA Vulkan, DX, OpenGL (NVIDIA GPUs only)
Xcode GPU Debugger Apple Metal (macOS / iOS only)

Integration in This Project

The Vulkan backend integrates two layers of RenderDoc support:

  1. VK_EXT_debug_utils — debug labels and object names baked into the command buffer, providing readable annotations inside RenderDoc's event browser.
  2. RenderDoc In-Application API (renderdoc_app.h) — runtime detection of RenderDoc, enabling F12 manual capture and --capture <frame> CLI auto-capture from within the application.

Per-Frame Command Buffer Structure

A single frame records the following sequence into one VkCommandBuffer:

# Event Debug Label Description
1 vkCmdResetQueryPool — Reset timestamp query slots for this frame
2 vkCmdWriteTimestamp (TOP_OF_PIPE) — T0: frame start
3 vkCmdBindPipeline (COMPUTE) Particle Compute (green) Bind compute pipeline
4 vkCmdBindDescriptorSets Bind SSBO descriptor (set 0, binding 0)
5 vkCmdPushConstants Push deltaTime (float) + damping (float)
6 vkCmdDispatch(N/256, 1, 1) Dispatch compute workgroups
7 vkCmdWriteTimestamp (COMPUTE_SHADER) — T1: compute end
8 vkCmdPipelineBarrier SSBO Barrier (yellow) SHADER_WRITE → VERTEX_ATTRIBUTE_READ
9 vkCmdWriteTimestamp (TOP_OF_PIPE) — T2: render start
10 vkCmdBeginRenderPass Particle Render (blue) Clear to (0.04, 0.08, 0.14, 1.0)
11 vkCmdBindPipeline (GRAPHICS) Bind graphics pipeline
12 vkCmdBindVertexBuffers Bind particle SSBO as VBO
13 vkCmdDraw(particleCount, 1, 0, 0) Draw all particles as POINT_LIST
14 vkCmdEndRenderPass Finish render pass
15 vkCmdWriteTimestamp (COLOR_OUTPUT) — T3: render end

Named Vulkan Objects

Object VK_EXT_debug_utils Name Purpose
particleBuffer_ Particle SSBO Interleaved position + velocity + colour buffer, used as both compute SSBO and vertex VBO
computePipeline_ Compute Pipeline Particle physics update (gravity, damping, boundary reflection)
graphicsPipeline_ Graphics Pipeline Point-sprite rendering with additive blending
renderPass_ Main Render Pass Single subpass, colour-only attachment

What to Inspect in RenderDoc

1. SSBO Data (Particle Buffer)

  • Select the vkCmdDispatch event → Pipeline State → Compute Shader → Descriptor Set 0 → Binding 0.
  • View buffer contents with custom struct format: float2 pos; float2 vel; float4 col;
  • Compare particle positions before and after the dispatch using RenderDoc's timeline — values should change by velocity × deltaTime.
  • Verify no NaN or out-of-bounds values.

2. Pipeline State

Stage Key Settings
Compute local_size = (256, 1, 1), 1 SSBO, 2 push constants
Vertex Input Binding 0 stride = 32 bytes; attr 0 = position (R32G32_SFLOAT, offset 0), attr 1 = colour (R32G32B32A32_SFLOAT, offset 16)
Rasteriser POINT_LIST topology, point size = 2.0 (set in vertex shader)
Blend Additive: srcColourBlendFactor = SRC_ALPHA, dstColourBlendFactor = ONE

3. Barrier Correctness

The single vkCmdPipelineBarrier between compute and render ensures:

  • Source: VK_PIPELINE_STAGE_COMPUTE_SHADER_BIT / VK_ACCESS_SHADER_WRITE_BIT
  • Destination: VK_PIPELINE_STAGE_VERTEX_INPUT_BIT / VK_ACCESS_VERTEX_ATTRIBUTE_READ_BIT
  • Scope: single buffer (Particle SSBO), VK_WHOLE_SIZE.

No additional implicit barriers should be inserted by the driver. If any appear in RenderDoc's event list, they indicate suboptimal synchronisation.

4. GPU Timing Cross-Validation

Compare RenderDoc's built-in per-event timing against the application's own vkCmdWriteTimestamp results:

Metric App Timestamps (ms) Notes
Compute dispatch 0.033 vkCmdWriteTimestamp T0 → T1
Render pass 0.408 vkCmdWriteTimestamp T2 → T3 (includes swapchain semaphore wait — see Section 6)
Total GPU time 0.446 T0 → T3 (RX 9070 XT, 1M particles)

The Chrome JSON from renderdoccmd convert records CPU-side API call durations (nanosecond resolution). For GPU-side per-event timing, open the .rdc capture in RenderDoc GUI → Window → Performance Counter Viewer. Deviation between app timestamps and RenderDoc GPU counters should be < 5 %. Larger discrepancies may indicate RenderDoc interception overhead or single-frame vs multi-frame averaging.

5. Potential Optimisations (identified via RenderDoc)

  • Vulkan 1.3 barrier upgrade: Replace VERTEX_INPUT_BIT with the more precise VK_PIPELINE_STAGE_2_VERTEX_ATTRIBUTE_INPUT_BIT to reduce stall scope.
  • Indirect dispatch: Replace hardcoded vkCmdDispatch(N/256, 1, 1) with vkCmdDispatchIndirect to allow GPU-driven workload sizing.
  • Dynamic point size: Move the hardcoded gl_PointSize = 2.0 to a push constant for runtime adjustment without pipeline recreation.

How to Capture

Build first with scripts\build-windows.ps1 (or CMake). Prefer the preset CLI path below; build\Release\... remains valid only if you configured -B build.

# Option A — GUI: launch from RenderDoc, press F12 during rendering
#   Executable: out\build\windows-x64-release\Release\gpu_benchmark.exe
#   (or gui\x64\Release\gpu_benchmark.exe after build-windows.ps1)
#   Working dir: same folder as the exe (shaders are adjacent)

# Option B — CLI (Windows)
& "C:\Program Files\RenderDoc\renderdoccmd.exe" capture `
    .\out\build\windows-x64-release\Release\gpu_benchmark.exe --backend vulkan --benchmark 200

# Option C — Auto-capture frame 50 (must be launched via RenderDoc)
.\out\build\windows-x64-release\Release\gpu_benchmark.exe --backend vulkan --benchmark 200 --capture 50
# Linux
renderdoccmd capture ./build/gpu_benchmark --backend vulkan --benchmark 200

Automated Cross-API Capture Analysis

The Full Analysis workflow (menu option 5/6) automatically captures one frame per API backend via the RenderDoc In-Application API, converts each .rdc to Chrome JSON using renderdoccmd convert, and runs rdoc_analyse.py to produce a structural comparison across all four Windows backends.

Per-Frame Event Count Comparison (RX 9070 XT, 1M particles, Medium):

API Total Events Frame Events Dispatches Draw Calls Barriers
Vulkan 1.2 77 28 1 1 1
DirectX 12 74 30 1 1 5
DirectX 11 86 26 1 1 0
OpenGL 4.3 90 17 1 1 1

Key Observations

  1. DX12 requires 5 resource barriers to accomplish what Vulkan and OpenGL each handle with a single barrier. This reflects DX12's finer-grained resource state tracking — each buffer/texture transition (e.g. UNORDERED_ACCESS → VERTEX_AND_CONSTANT_BUFFER, RENDER_TARGET → PRESENT) is an explicit barrier. Vulkan batches the same transitions into one vkCmdPipelineBarrier call with multiple memory barriers.

  2. DX11 has zero explicit barriers. The DX11 driver silently inserts all necessary synchronisation on behalf of the application. This is the key trade-off of implicit APIs: simpler code at the cost of opaque scheduling decisions that profiling tools like RenderDoc cannot surface.

  3. OpenGL's frame has the fewest events (17 frame events) because the OpenGL driver consolidates many state changes into fewer internal calls. The single glMemoryBarrier(GL_VERTEX_ATTRIB_ARRAY_BARRIER_BIT) maps directly to the same compute → vertex synchronisation as Vulkan's vkCmdPipelineBarrier.

  4. Vulkan debug labels are visible in captures — 6 vkCmdBeginDebugUtilsLabelEXT / vkCmdEndDebugUtilsLabelEXT calls (3 pairs: "Particle Compute", "SSBO Barrier", "Particle Render") and 4 vkCmdWriteTimestamp calls confirm the profiling instrumentation is correctly captured.

  5. DX12 has the most frame events (30) despite having fewer total events. This is because explicit APIs expose more per-frame work: command allocator reset, command list recording, root signature binding, and explicit resource transitions that implicit APIs hide from the application.

Synchronisation Model Comparison

Vulkan DX12 DX11 OpenGL
Barrier model vkCmdPipelineBarrier — stage + access masks ResourceBarrier — state transitions Implicit (driver-managed) glMemoryBarrier — bit flags
Barriers per frame 1 5 0 1
Who manages sync? Application Application Driver Application (coarse)
Profiler visibility Full Full Hidden Partial

The full per-event command sequences for all APIs are in docs/rdoc_comparison.md, auto-generated by scripts/rdoc_analyse.py.

See docs/renderdoc-analysis.md for the detailed Vulkan analysis template and docs/renderdoc-capture-guide.md for step-by-step capture instructions.


4. Python Benchmark Tooling

Automated data analysis scripts in the scripts/ directory.

Script Purpose
plot_results.py Read results.json and generate 4 charts: FPS by GPU × API, GPU time breakdown, CPU overhead, particle-count scaling
batch_benchmark.py Iterate over all GPU × API × particle-count combinations, invoke the benchmark executable, and collect results automatically
export_report.py Export results as markdown tables or a standalone sortable HTML report with dark theme
compare_3dmark.py Cross-validate against 3DMark: normalised bar chart, R² correlation scatter plot, auto-import from .3dmark-result files

All scripts read from ~/.gpu_bench/results.json (the application's auto-saved benchmark results). Charts use a dark colour scheme with API-specific colours (Vulkan = red, DX12 = blue, DX11 = green, OpenGL = orange, Metal = purple).


Part 2 — Primary GPUs: RX 9070 XT & RX 6900 XT

5. RX 9070 XT and RX 6900 XT — Cross-API, Particle Scaling, and Headless Analysis

AMD's first RDNA 4 discrete GPU, tested across three scenarios: standard 1M windowed, maximum 16M windowed, and headless compute. The 9070 XT provides a unique lens into swapchain throttling behaviour due to its very fast compute throughput relative to presentation overhead.

Test Hardware

Component Specification
CPU AMD Ryzen 5 7600 6-Core Processor
GPU AMD Radeon RX 9070 XT (RDNA 4, 64 CU, 16 GB GDDR6, 640 GB/s)
Driver AMD Adrenalin 26.3.1 (LLPC)
OS Windows 11 (NT 10.0.26200)
Resolution 1280 × 720
V-Sync OFF
Memory Mode Device-local

5a. Cross-API Comparison — 1M Particles (Medium), Windowed

RX 9070 XT:

# API Avg FPS Compute (ms) Render (ms) Total GPU (ms) GPU Util Bottleneck
1 DX11 1,773.7 0.047 0.451 0.542 100% GPU-bound
2 Vulkan 1,750.6 0.033 0.408 0.446 80% Balanced
3 DX12 1,608.7 0.034 0.399 0.434 70% Balanced
4 OpenGL 253.4 2.612 0.792 3.658 90% GPU-bound

RX 6900 XT:

# API Avg FPS Compute (ms) Render (ms) Total GPU (ms) GPU Util Bottleneck
1 DX11 4,067.9 0.047 0.140 0.226 90% GPU-bound
2 DX12 3,518.0 0.049 0.112 0.162 60% Balanced
3 Vulkan 2,884.6 0.063 0.133 0.196 60% Balanced
4 OpenGL 228.6 2.742 1.151 4.066 90% GPU-bound

Key observations:

  • At 1M particles, the 6900 XT achieves higher FPS (4,068 DX11 vs 1,774 on 9070 XT). This is again a CU-count advantage in this presentation-limited scenario — more CUs finish the trivial compute+render faster, leaving more frame budget for swapchain overhead.
  • DX11 and Vulkan are nearly tied on the 9070 XT at ~1750 FPS. Unlike on RTX 5090 where DX11 dominates (8955 FPS vs 3611 Vulkan), the gap is much smaller on AMD — the AMD DX11 driver does not have the same micro-optimisation advantage as NVIDIA's.
  • DX11 leads on both GPUs — on the 6900 XT (4,068 FPS), DX11 is 16% faster than DX12 (3,518 FPS) and 41% faster than Vulkan (2,885 FPS). At these extreme frame rates, DX11's lower per-frame CPU overhead dominates.
  • OpenGL is severely penalised on both GPUs — the AMD OpenGL compute overhead issue (Section 16) causes ~2.6–2.8 ms compute time vs 0.03–0.05 ms on Vulkan, a ~70× penalty.
  • All APIs show inflated render times (0.1–0.8 ms) due to swapchain semaphore wait pollution (detailed in Section 6). The actual render work is ~0.04 ms.

5b. Cross-API Comparison — 16M Particles (Ultra), Windowed

RX 9070 XT:

# API Avg FPS Compute (ms) Render (ms) Total GPU (ms) GPU Util Bottleneck
1 Vulkan 110.9 1.963 6.440 8.418 90% GPU-bound
2 DX12 105.7 1.940 6.126 8.069 90% GPU-bound
3 DX11 93.5 1.908 6.152 9.904 90% GPU-bound
4 OpenGL 15.7 47.959 10.073 58.473 90% GPU-bound

RX 6900 XT:

# API Avg FPS Compute (ms) Render (ms) Total GPU (ms) GPU Util Bottleneck
1 Vulkan 218.4 2.145 2.079 4.224 90% GPU-bound
2 DX12 205.4 2.113 1.851 4.028 90% GPU-bound
3 DX11 133.9 2.121 1.966 6.664 90% GPU-bound
4 OpenGL 14.3 45.609 19.009 64.759 90% GPU-bound

Side-by-side 16M comparison (9070 XT vs 6900 XT):

API 9070 XT FPS 6900 XT FPS 9070 XT Compute 6900 XT Compute 9070 XT Render 6900 XT Render
Vulkan 110.9 218.4 1.963 ms 2.145 ms 6.440 ms 2.079 ms
DX12 105.7 205.4 1.940 ms 2.113 ms 6.126 ms 1.851 ms
DX11 93.5 133.9 1.908 ms 2.121 ms 6.152 ms 1.966 ms
OpenGL 15.7 14.3 47.959 ms 45.609 ms 10.073 ms 19.009 ms

An interesting result: the RX 6900 XT outperforms the newer RX 9070 XT at 16M particles (218 vs 111 FPS on Vulkan), despite the 9070 XT scoring 1.4× higher in Time Spy and 1.7× higher in Steel Nomad — and being faster in both pure compute (headless, Section 5c) and 1M windowed. The compute times are nearly identical (~2.0 ms); the difference is entirely in the render pass: 6.4 ms on the 9070 XT vs 2.1 ms on the 6900 XT.

This is a workload-specific result that highlights what this microbenchmark actually measures. Rendering 16M individual point primitives is an extreme stress test for primitive throughput — the ability to set up and rasterise millions of tiny primitives per frame. The 6900 XT has 80 CUs with 1.25× more rasterisation hardware than the 9070 XT's 64 CUs, but the render time gap (3.1×) is far larger than the CU ratio alone would suggest.

Infinity Cache plays a significant role here. AMD has progressively reduced Infinity Cache capacity across RDNA generations — 128 MB (RDNA 2) → 96 MB (RDNA 3, e.g. RX 7900 XTX) → 64 MB (RDNA 4) — trading raw capacity for improved per-MB efficiency. However, in this 16M-particle scenario, the 6900 XT's 2× larger cache (128 MB vs 64 MB) is a clear advantage: 16M point primitives generate heavy vertex fetch and rasterisation traffic, and the larger cache keeps more of this data on-die, reducing round-trips to VRAM. The 9070 XT compensates with higher memory bandwidth (640 vs 512 GB/s), but bandwidth cannot fully offset the latency penalty when the working set exceeds the cache. This mirrors a well-known pattern in CPUs — AMD's Ryzen 9800X3D with 3D V-Cache (96 MB L3) dramatically outperforms the standard 9700X (32 MB L3) in cache-sensitive gaming workloads, despite identical core counts and clocks. Whether on a GPU or CPU, when the working set fits in a larger cache, the raw bandwidth of the smaller-cache part cannot compensate for the hit-rate advantage.

Beyond cache, the 6900 XT also benefits from a more mature RDNA 2 driver pipeline for point primitive rendering. Standard game rendering uses far fewer, larger triangles with complex shading — a scenario where the 9070 XT's architectural improvements (higher clocks, ray tracing hardware, improved schedulers) deliver the performance uplift reflected in 3DMark. The lesson is that no single benchmark captures all aspects of GPU performance; this microbenchmark specifically targets compute + primitive throughput, which produces a different ranking than rasterisation-focused benchmarks with complex geometry.

16× particle scaling analysis (1M → 16M) — RX 9070 XT:

API 1M Compute 16M Compute Scaling (expected 16×) 1M FPS 16M FPS FPS Ratio
Vulkan 0.033 1.963 59.5× 1,750.6 110.9 15.8×
DX12 0.034 1.940 57.1× 1,608.7 105.7 15.2×
DX11 0.047 1.908 40.6× 1,773.7 93.5 19.0×
OpenGL 2.612 47.959 18.4× 253.4 15.7 16.1×
  • Compute time scales super-linearly (57–60× for 16× particles on Vulkan/DX12). This is expected: at 1M particles the GPU is underutilised and the per-dispatch overhead dominates; at 16M particles the ALUs and memory bandwidth are fully saturated.
  • FPS scales roughly linearly (~16× reduction) because at 16M particles all APIs are GPU-bound — CPU overhead is negligible relative to GPU execution time.
  • DX11's compute–render gap widens: total GPU (9.904 ms) exceeds compute + render sum (1.908 + 6.152 = 8.060 ms) by 1.844 ms, suggesting pipeline synchronisation overhead similar to what was observed on the GTX 970 (Section 16).
  • OpenGL's compute time explodes to 47.959 ms — the AMD OpenGL dispatch overhead scales worse than linearly with particle count, making OpenGL completely impractical for high particle counts on AMD.
  • Render time dominates at 16M: ~6 ms for Vulkan/DX12/DX11 vs ~0.4 ms at 1M. This is real render work (16M point sprites), not semaphore wait — at 16M particles the GPU is genuinely busy rendering, unlike at 1M where the semaphore wait dominated.

5c. Headless Compute — 1M Particles

RX 9070 XT:

# API Avg FPS Compute (ms) Render (ms) Total GPU (ms) GPU Util Bottleneck
1 DX12 21,354.0 0.034 0.0 0.035 70% Balanced
2 Vulkan 21,259.6 0.034 0.0 0.034 70% Balanced
3 OpenGL 20,298.3 0.034 0.0 0.034 70% Balanced
4 DX11 16,937.8 0.034 0.0 0.034 60% Balanced

RX 6900 XT:

# API Avg FPS Compute (ms) Render (ms) Total GPU (ms) GPU Util Bottleneck
1 DX12 15,949.6 0.048 0.0 0.049 70% Balanced
2 Vulkan 15,241.5 0.055 0.0 0.055 70% Balanced
3 DX11 12,274.7 0.047 0.0 0.052 60% Balanced
4 OpenGL 341.2 2.750 0.0 2.750 90% GPU-bound

Headless cross-GPU comparison (1M particles):

API 9070 XT FPS 6900 XT FPS 9070 XT / 6900 XT
Vulkan 21,260 15,242 1.39×
DX12 21,354 15,950 1.34×
DX11 16,938 12,275 1.38×
OpenGL 20,298 341 59.5×

In headless mode, the 9070 XT is consistently ~1.37× faster than the 6900 XT in pure compute — this is the true generational improvement, free of presentation throttling and render pipeline differences. The 9070 XT achieves this with 64 CUs vs 80 CUs (80% of the CU count), meaning its per-CU compute efficiency is ~1.7× higher than RDNA 2. The 6900 XT's OpenGL headless result (341 FPS vs 15K+ on other APIs) confirms the AMD OpenGL compute overhead persists even without a window.

Why Infinity Cache no longer matters here. In Section 5b, the 6900 XT's 128 MB Infinity Cache was identified as a key factor in its 16M-particle render advantage over the 9070 XT (64 MB). In headless mode, that advantage disappears entirely — the 9070 XT wins by 1.37×. The reason is a fundamental difference in memory access patterns: the compute shader performs a streaming update (each particle reads its own position/velocity, updates, and writes back), where data is touched once and not reused — cache size is irrelevant, and raw bandwidth (640 vs 512 GB/s) and ALU throughput determine performance. By contrast, the 16M-particle render pass involves vertex fetch, primitive assembly, rasterisation, and depth testing across the same memory regions, creating repeated, overlapping accesses where a larger cache dramatically improves hit rates. This confirms that Infinity Cache is a render-path advantage in cache-sensitive workloads, not a universal compute advantage. Paradoxically, headless compute results are a better proxy for traditional gaming rasterisation performance at this GPU tier than the 16M-particle windowed render test. Real game rendering uses thousands of large, shaded triangles — a workload dominated by ALU throughput, bandwidth, and clock speed rather than raw primitive count. The headless 1.37× advantage for the 9070 XT aligns closely with its 1.4× Time Spy and 1.7× Steel Nomad leads, while the 16M-particle render test — with its extreme small-primitive throughput stress — is an outlier that specifically punishes smaller caches and fewer fixed-function rasterisation units. In other words, the windowed 16M test tells us something real about the hardware, but it is the headless result that better predicts how these GPUs rank in the workloads most users care about.

RX 9070 XT — Windowed vs Headless comparison:

API Windowed FPS Headless FPS Speedup Windowed Compute Headless Compute
Vulkan 1,750.6 21,259.6 12.1× 0.033 ms 0.034 ms
DX12 1,608.7 21,354.0 13.3× 0.034 ms 0.034 ms
DX11 1,773.7 16,937.8 9.5× 0.047 ms 0.034 ms
OpenGL 253.4 20,298.3 80.1× 2.612 ms 0.034 ms
  • All four APIs converge to identical compute time (0.034 ms) on the 9070 XT in headless mode, proving the GPU compute hardware is equivalent regardless of API.
  • OpenGL's 80× speedup is the most dramatic — headless mode bypasses the AMD OpenGL compute dispatch overhead entirely, as the dispatch path is simpler without a rendering context.
  • DX11 compute drops from 0.047 to 0.034 ms in headless, suggesting the DX11 driver's implicit state management adds ~0.013 ms overhead even to compute dispatch when a swapchain is present.
  • The remaining FPS differences (21K vs 17K for DX11) reflect pure CPU-side overhead differences between APIs.

5d. Flights Test — 1M Particles, Windowed (2 vs 3 Frames-in-Flight)

What is "Frames-in-Flight"? In modern graphics APIs, the CPU does not wait for the GPU to finish one frame before starting the next. Instead, the CPU can prepare N frames ahead while the GPU is still rendering earlier ones — these N in-progress frames are called "frames-in-flight" (also known as "buffered frames" or "frame overlap"). This is essentially a render queue — the number of frames simultaneously in-progress in the CPU–GPU pipeline, analogous to the "work-in-progress" slots on a factory assembly line. With 2 flights, the CPU can be building frame N+1 while the GPU renders frame N; with 3 flights, the CPU can be up to 2 frames ahead. A deeper queue improves throughput (the pipeline is less likely to stall), but increases input latency (the displayed frame was prepared further in the past). This is the same concept exposed by NVIDIA's "Maximum Pre-Rendered Frames" setting (now called "Low Latency Mode") and targeted by latency-reduction technologies like NVIDIA Reflex and AMD Anti-Lag, which dynamically shorten the render queue to minimise input-to-display delay at the cost of some throughput. This test compares 2 vs 3 frames-in-flight to measure the throughput impact on each API.

Relationship with V-Sync. Frames-in-flight and V-Sync are related but distinct concepts. V-Sync locks presentation to the display refresh rate to prevent tearing; frames-in-flight controls how far the CPU can work ahead of the GPU regardless of V-Sync. The two interact most visibly when V-Sync is ON: on a 60 Hz display, 2 flights means the CPU leads by up to 1 frame (~17 ms extra latency), while 3 flights ("triple buffering") allows up to 2 frames ahead (~33 ms extra latency). With V-Sync OFF — the condition used in all tests in this report — flights purely affect CPU–GPU pipeline overlap without any refresh-rate constraint, which is why DX12 can gain +22% throughput simply by moving from 2 to 3 flights: the CPU submits work earlier instead of stalling on a busy swapchain image.

API Flights=2 FPS Flights=3 FPS Change Flights=2 Render Flights=3 Render
Vulkan 1,750.6 1,736.9 −0.8% 0.408 ms 0.409 ms
DX12 1,608.7 1,964.5 +22.1% 0.399 ms 0.400 ms
DX11 1,773.7 1,956.3 +10.3% 0.451 ms 0.399 ms
OpenGL 253.4 256.1 +1.1% 0.792 ms 0.781 ms
  • DX12 benefits most from an extra frame-in-flight (+22%), suggesting its command pipeline can overlap more work with 3 buffers.
  • Vulkan shows no improvement — its presentation engine already manages buffering efficiently at 2 frames.
  • Render times remain unchanged across both flight counts, confirming that swapchain semaphore wait pollution is not reduced by adding more swapchain images (as discussed in Section 6d).
  • Why this benchmark defaults to 2 flights. For GPU benchmarking, fewer frames-in-flight is generally preferable: it minimises CPU-side pipeline overlap so that the measured frame time more closely reflects actual GPU execution cost. With 3 flights, the CPU can "hide" stalls by working further ahead, which inflates FPS without the GPU doing any more work per unit time — this is a CPU-side throughput optimisation, not a GPU performance improvement. The DX12 +22% gain from 3 flights demonstrates exactly this effect: the GPU render time is unchanged (0.399 → 0.400 ms), meaning the GPU is doing identical work; the extra FPS comes entirely from the CPU submitting commands more efficiently. For a benchmark designed to measure GPU compute and render performance, 2 flights provides a cleaner signal with less CPU-side noise. This section tests 3 flights to quantify the presentation overhead, not to suggest it as a better default.

5e. RX 9070 XT vs Other GPUs — 1M Particles, Vulkan

GPU Architecture Compute (ms) Render (ms) Total GPU (ms) FPS
RTX 5090 Blackwell (170 SM) 0.025 0.064 0.090 3,611
RX 9070 XT RDNA 4 (64 CU) 0.033 0.408 0.446 1,751
RX 6900 XT RDNA 2 (80 CU) 0.063 0.134 0.197 2,866
RX 6600 XT RDNA 2 (32 CU) 0.270 0.379 0.649 1,239
Vega FE GCN 5 (64 CU) 0.219 0.275 0.495 1,370
RX 580 GCN 4 (36 CU) 0.362 0.702 1.070 783

Compute performance ranking (lower is better):

GPU Compute (ms) vs RX 9070 XT Per-CU Efficiency vs 9070 XT
RTX 5090 0.025 0.76× N/A (different arch)
RX 9070 XT 0.033 1.00× 1.00×
RX 6900 XT 0.063 1.91× 0.42× (80 CU → per-CU: 1.91 × 64/80 = 1.53× slower)
Vega FE 0.219 6.64× 0.15× (64 CU → per-CU: 6.64× slower)
RX 6600 XT 0.270 8.18× 0.24× (32 CU → per-CU: 8.18 × 64/32 = 4.1× slower)
RX 580 0.362 10.97× 0.16× (36 CU → per-CU: 10.97 × 64/36 = 6.2× slower)

The 9070 XT is 8.2× faster overall than the 6600 XT in compute. Since the 9070 XT has 2× the CU count (64 vs 32), the per-CU efficiency improvement is ~4.1×, which is still a massive generational leap from RDNA 2 to RDNA 4:

  • Higher clock speed (2,970 MHz vs 2,589 MHz) accounts for ~1.15×
  • 2× the CU count accounts for another ~2×
  • The remaining ~1.8× improvement comes from architectural changes: improved compute scheduler, better cache hierarchy, wider memory interface utilisation (640 vs 256 GB/s), and mature RDNA 4 driver code generation

The 9070 XT also outperforms the 80-CU RX 6900 XT in compute (0.033 vs 0.063 ms) despite having 80% of its CU count, making it the fastest AMD compute GPU tested in this benchmark.


6. Swapchain Throttling, Timestamp Pollution, and Headless Compute

6a. The Problem: Why Fast GPUs Appear Slower Than Expected

During windowed benchmark runs, the RX 9070 XT exhibited an unexpected anomaly: despite computing 2× faster than the RX 6900 XT, its reported render time was higher, resulting in lower-than-expected total GPU time efficiency.

GPU Compute (ms) Render (ms) Total GPU (ms) FPS
RTX 5090 0.025 0.064 0.090 3,611
RX 6900 XT 0.063 0.134 0.197 2,866
RX 9070 XT 0.033 0.408 0.441 1,981

The 9070 XT's render time (0.408 ms) is 3× that of the 6900 XT (0.134 ms), despite rendering the same single draw call of point sprites. Investigation revealed this is not a GPU performance issue but a measurement artefact caused by swapchain semaphore wait pollution in timestamps.

6b. Root Cause: Semaphore Wait in Vulkan Timestamps

In the Vulkan backend, timestamp T3 is written at VK_PIPELINE_STAGE_COLOR_ATTACHMENT_OUTPUT_BIT. This pipeline stage does not begin until the presentation engine releases a swapchain image — signalled via the imageAvailableSemaphore acquired from vkAcquireNextImageKHR.

T2 (TOP_OF_PIPE)  ──→  Vertex Shader  ──→  Rasterisation  ──→  Fragment Shader
                                                                      │
                                                        ┌─────────────┘
                                                        ▼
                                             COLOR_ATTACHMENT_OUTPUT
                                             ┌──────────────────────┐
                                             │  Wait for semaphore  │ ← swapchain image availability
                                             │  (variable delay)    │
                                             │  Actual pixel write  │
                                             └──────────────────────┘
                                                        │
                                                        ▼
                                                   T3 (timestamp)

The render time (T3 − T2) therefore measures: actual render work + semaphore wait for swapchain image. On fast GPUs that finish compute + render in under 1 ms, the GPU spends most of its time idle, waiting for the presentation engine to recycle a swapchain image from the previous frame's vkQueuePresentKHR.

6c. Why Different GPUs Are Affected Differently

GPU Situation Effect on Render Timestamp
RTX 5090 CPU-bound (30% GPU util). GPU finishes early, waits for next CPU submission. Semaphore wait is hidden within CPU stall. Minimal pollution — GPU idle time is between frames, not during render stage
RX 6900 XT Balanced. Compute takes long enough (0.063 ms) that the presentation engine has time to release the swapchain image before render stage begins. Low pollution — semaphore is usually already signalled
RX 9070 XT Compute finishes very fast (0.033 ms), reaches COLOR_ATTACHMENT_OUTPUT before the presentation engine has released the image from the previous present. GPU stalls waiting for semaphore. High pollution — ~0.37 ms of the 0.408 ms "render time" is semaphore wait

The 9070 XT's actual render work is approximately 0.04 ms (similar to other GPUs), but the timestamp reports 0.408 ms because it includes the swapchain image wait.

Note that at 1M particles, the render timestamp difference is primarily a semaphore wait artefact, not a real GPU performance gap. However, at higher particle counts (16M), the render time difference becomes real — and Infinity Cache becomes a significant factor: the 6900 XT's 128 MB cache vs the 9070 XT's 64 MB provides a substantial hit-rate advantage for primitive-heavy rendering (see Section 5b for full analysis).

6d. Swapchain BufferCount vs VSync

Two frequently confused concepts that both affect frame pacing:

Swapchain BufferCount VSync
What it controls How many swapchain images exist in the pool When completed frames are shown on the display
Set by VkSwapchainCreateInfoKHR::minImageCount (Vulkan), DXGI_SWAP_CHAIN_DESC::BufferCount (DX11/12) glfwSwapInterval() (OpenGL), Present(syncInterval) (DX), VK_PRESENT_MODE_* (Vulkan)
Effect on FPS More buffers → less GPU idle time waiting for image availability, but diminishing returns beyond 3 VSync ON → FPS capped to display refresh rate; VSync OFF → uncapped
Effect on input lag More buffers → higher input lag (more pre-rendered frames queued) VSync ON → adds up to one frame of latency

Increasing BufferCount from 2 to 3 was tested (via --flights 3) but showed no meaningful improvement in timestamp pollution. The semaphore wait is inherent to the Vulkan presentation model — the timestamp at COLOR_ATTACHMENT_OUTPUT_BIT will always include the wait, regardless of how many images are in the pool, because the GPU must still wait for at least one image to become available from the presentation engine.

6e. Headless Compute Mode — Eliminating Presentation Overhead

To measure pure GPU compute performance without any swapchain, rendering, or presentation interference, a headless compute mode (--headless) was implemented:

Component Windowed Mode Headless Mode
Window GLFW visible window No window (OpenGL: hidden window for context)
Swapchain Created, images acquired/presented Not created
Render pass Full vertex + fragment pipeline Skipped entirely
Present vkQueuePresentKHR / Present() / glfwSwapBuffers Skipped
Timestamps T0–T3 (compute + render) T0–T1 (compute only), T2=T3 mirrored
GPU utilisation Limited by presentation engine Limited only by compute throughput

Headless Results — RX 9070 XT, 1M Particles

API FPS Compute (ms) GPU Util
Vulkan 21,260 0.034 70%
DX12 21,354 0.034 70%
DX11 16,938 0.034 60%
OpenGL 20,298 0.034 70%

All four APIs converge to nearly identical compute times (0.034 ms), confirming that the GPU-side compute workload is equivalent across APIs. The FPS difference reflects only CPU-side overhead — DX11 is slightly slower due to its implicit driver model requiring more CPU work per frame without a Present() call to batch around.

Compare with windowed mode:

API Windowed FPS Headless FPS Speedup
Vulkan 1,981 21,260 10.7×
DX12 2,100 21,354 10.2×
DX11 3,500 16,938 4.8×
OpenGL 1,800 20,298 11.3×

The 10× speedup confirms that windowed mode performance is dominated by presentation overhead, not by compute or render workload.

6f. API-Specific Headless Implementation Challenges

DX11: No Frame Boundary Without Present()

DX11's implicit driver model uses Present() as an implicit frame boundary for command batching and timestamp query resolution. Without it:

  • Problem 1: CollectTimestampResults() used Sleep(1) retries waiting for query resolution. Without Present(), queries never resolved promptly, causing each frame to take 4+ ms (Sleep granularity).
  • Fix: Removed Sleep in headless mode; spin-wait only.
  • Problem 2: Even with spin-wait, timestamp values were occasionally garbage (e.g., 805534675707 ms) because DX11 lacks proper frame boundaries without Present().
  • Fix: Added context_->Flush() after compute dispatch to force command submission, plus sanity filter discarding timestamps > 1000 ms. Approximately 3–4% of frames produce garbage values and are discarded (e.g., 212531/220193 valid samples).

OpenGL: AMD Driver Requires Explicit Flush for Hidden Windows

OpenGL requires a window (even hidden) to create a GL context. On AMD drivers, glFlush() alone is insufficient to process commands for hidden windows — the driver does not actively schedule GPU work without a visible surface.

  • Attempt 1: glFlush() only → timestamps never resolve (0 valid samples).
  • Attempt 2: glFinish() every frame → timestamps work, but FPS drops from 22,000 to 8,000 (CPU stalls waiting for GPU).
  • Final solution: glFinish() every 16th frame (forces command processing) + glFenceSync + glFlush() on other frames (non-blocking). This achieves 20,298 FPS with ~25% timestamp sample rate (65985/263889 valid samples).

Vulkan and DX12: Clean Headless

Both explicit APIs handle headless cleanly:

  • Skip swapchain, render pass, and present calls
  • Compute dispatch + fence sync is sufficient
  • 100% timestamp sample rate, no workarounds needed

6g. Comparison with 3DMark Unlimited Mode

3DMark offers an Unlimited mode that removes VSync and frame rate caps. This is often confused with headless compute, but they are fundamentally different:

3DMark Unlimited This Benchmark Headless
Rendering Full offscreen rendering (all geometry, textures, post-FX) No rendering — compute dispatch only
Target Offscreen render target (no swapchain present) No render target at all
Measures Combined compute + render + post-processing GPU throughput, uncapped Pure compute shader throughput
Presentation Skipped (no VSync, no Present) Skipped
Use case Cross-device comparison without display refresh rate bias Isolating compute performance from presentation overhead
Analogy Running the full game engine but rendering to a texture instead of screen Running only the physics engine with no rendering at all

3DMark Unlimited is equivalent to rendering to an offscreen framebuffer — the full GPU pipeline (vertex → rasterisation → fragment → post-processing) executes, but the final present/flip is skipped. This benchmark's headless mode is more aggressive: it eliminates the entire graphics pipeline, measuring only the compute dispatch that updates particle positions.

If a 3DMark Unlimited-style mode were added to this benchmark, it would involve:

  1. Creating an offscreen framebuffer (VkFramebuffer / ID3D11RenderTargetView / FBO)
  2. Running the full compute + render pipeline to that framebuffer
  3. Skipping only vkQueuePresentKHR / Present() / glfwSwapBuffers
  4. Timestamp T3 would measure actual render completion without semaphore wait pollution

This would provide a middle ground between windowed (presentation-throttled) and headless (compute-only) modes, and would be the most direct comparison point with 3DMark Unlimited scores.


Part 3 — AMD GPU Generational Analysis

7. AMD GPU Generational Analysis — TeraScale 2 → GCN → RDNA 2 → RDNA 4

Cross-generational compute shader performance comparison across eight AMD GPUs spanning five architectures and 16 years of hardware evolution (2009–2025). All results collected with this project's particle simulation benchmark.

Test Hardware

GPU Architecture Generation CUs / SPs Core Clock FP32 TFLOPS Memory Bandwidth Platform API Coverage
HD 5770 TeraScale 2 2009 800 SPs (VLIW5) 850 MHz ~1.36 1 GB GDDR5 76.8 GB/s Windows DX11, OpenGL
FirePro D700 GCN 1.0 (Tahiti) 2013 2048 SPs 850 MHz ~3.5 6 GB GDDR5 264 GB/s Windows Vulkan, DX12, DX11, OpenGL
RX 580 GCN 4 (Polaris) 2017 36 CUs 1,340 MHz ~6.2 8 GB GDDR5 256 GB/s Windows Vulkan, DX12, DX11, OpenGL
Vega Frontier Edition GCN 5 (Vega) 2017 64 CUs 1,600 MHz ~13.1 16 GB HBM2 483 GB/s Windows Vulkan, DX12, DX11, OpenGL
RX 6600 XT RDNA 2 2021 32 CUs 2,589 MHz ~10.6 8 GB GDDR6 256 GB/s Windows Vulkan, DX12, DX11, OpenGL
RX 6900 XT RDNA 2 2020 80 CUs 2,250 MHz ~23.0 16 GB GDDR6 512 GB/s Windows Vulkan, DX12, DX11, OpenGL
RX 9070 XT RDNA 4 2025 64 CUs 2,970 MHz ~48.7 † 16 GB GDDR6 640 GB/s Windows Vulkan, DX12, DX11, OpenGL
Ryzen 9800X3D iGPU RDNA 2 2024 2 CUs 2,200 MHz ~0.56 Shared DDR5 ~83 GB/s Windows Vulkan, DX12, DX11, OpenGL

Note: FirePro D700 data is from Windows (Boot Camp). Vulkan, DX12, DX11, and OpenGL results are available for GCN 1.0. All GPUs are tested on Windows. † RDNA 4 uses dual-issue FP32 (each SP executes 2 FP32 ops/clock); traditional single-issue calculation yields ~24.3 TFLOPS.


7a. Raw Results — All AMD GPUs, 1M Particles (Medium)

Best API per GPU (highest FPS):

# GPU Architecture Best API Avg FPS Compute (ms) Render (ms) Total GPU (ms) Bottleneck
1 RX 9070 XT RDNA 4 (64 CU) DX11 1,773.7 0.047 0.451 0.542 GPU-bound
2 RX 6900 XT RDNA 2 (80 CU) DX11 4,067.9 0.047 0.140 0.226 GPU-bound
3 RX 6600 XT RDNA 2 (32 CU) DX12 1,834.4 0.190 0.215 0.406 Balanced
3 Vega FE GCN 5 (64 CU) DX12 1,715.5 0.219 0.229 0.452 Balanced
4 RX 580 GCN 4 (36 CU) DX12 912.2 0.362 0.557 0.930 GPU-bound
5 FirePro D700 GCN 1.0 Vulkan 554.8 0.589 0.881 1.473 GPU-bound
6 Zen4/5 iGPU (2 CU) RDNA 2 DX12 324.0 1.480 1.472 2.953 GPU-bound
7 HD 5770 TeraScale 2 OpenGL 188.2 1.794 3.018 4.818 GPU-bound
8 WARP on 9800X3D¹ Software DX12 86.6 1.034 10.024 11.059 Software
9 WARP on 7600¹ Software DX12 62.2 1.810 13.369 15.181 Software

¹ WARP runs on the CPU, not a GPU. The two WARP entries show the same software renderer on different CPUs: the Ryzen 7 9800X3D (8-core, 96 MB L3 3D V-Cache) is 39% faster than the Ryzen 5 7600 (6-core, 32 MB L3), demonstrating how CPU core count, clock speed, and cache size directly affect software rendering performance.

All API results per GPU:

GPU Vulkan DX12 DX11 OpenGL Metal
RX 9070 XT 1,751 FPS 1,609 FPS 1,774 FPS 253 FPS N/A
RX 6900 XT 2,885 FPS 3,518 FPS 4,068 FPS 229 FPS N/A
RX 6600 XT 1,239 FPS 1,834 FPS 988 FPS 180 FPS N/A
Vega FE 1,370 FPS 1,716 FPS 1,436 FPS 158 FPS N/A
RX 580 783 FPS 912 FPS 755 FPS 42 FPS N/A
FirePro D700 555 FPS 516 FPS 527 FPS 525 FPS N/A
Zen4/5 iGPU (2 CU) 238 FPS 324 FPS 271 FPS 229 FPS N/A
HD 5770 N/A N/A 107 FPS 188 FPS N/A

7b. Per-CU Compute Efficiency

Per-CU Compute Efficiency

Normalise compute shader performance to per-CU throughput to isolate architectural efficiency from raw CU count.

GPU Architecture CUs Compute Time (ms) Per-CU Throughput (relative) Per-CU vs RX 580
HD 5770 TeraScale 2 ~10 equiv 1.794 (OpenGL) 0.0557 0.73×
FirePro D700 GCN 1.0 32 0.589 0.0531 0.69×
RX 580 GCN 4 36 0.362 0.0767 1.00×
Vega FE GCN 5 64 0.219 0.0713 0.93×
RX 6600 XT RDNA 2 32 0.270 0.1157 1.51×
RX 9070 XT RDNA 4 64 0.033 0.4735 6.17×
RX 6900 XT RDNA 2 80 0.063 0.1984 2.59×
Zen4/5 iGPU (2 CU) RDNA 2 2 1.257 0.3978 5.19×

HD 5770 CU equivalence: TeraScale 2 does not have CUs. 800 VLIW5 stream processors are roughly grouped into 10 SIMD engines. This mapping is approximate.

Analysis:

Per-CU efficiency improves modestly from GCN 1.0 (0.69x) through GCN 4 (1.00x) to GCN 5 (0.93x), with Vega FE slightly below RX 580 per-CU despite being a newer architecture -- likely because the 64 CUs are not fully utilised at 1M particles. The jump to RDNA 2 brings a 1.51x per-CU improvement on the RX 6600 XT, confirming the architectural efficiency gain from GCN to RDNA.

The standout result is RDNA 4: the RX 9070 XT achieves 6.17× the per-CU throughput of the RX 580, a significant generational leap. This is partly real architectural improvement (higher clocks, better scheduler, 640 GB/s bandwidth vs 256 GB/s) and partly because the 0.033 ms compute time is near the floor of per-dispatch overhead, meaning the GPU finishes so quickly that fixed overhead dominates less.

The iGPU (2 CU) shows 5.19x per-CU efficiency vs the RX 580, which is surprisingly high. With only 2 CUs, the workgroup scheduling overhead is minimal and the entire working set fits in cache, inflating per-CU throughput. The RX 6900 XT (2.59x) has lower per-CU than the iGPU despite being the same RDNA 2 architecture, confirming that CU scaling is sub-linear -- more CUs means more contention for memory bandwidth and cache.


7c. CU Scaling Within RDNA 2

Three RDNA 2 GPUs at vastly different CU counts allow direct measurement of how compute performance scales with CU count within the same architecture.

GPU CUs FPS Compute (ms) Scaling vs iGPU (2 CU) Ideal Scaling (CU ratio) Efficiency
Zen4/5 iGPU (2 CU) 2 238 1.257 1.00× 1.00× 100%
RX 6600 XT 32 1,239 0.270 4.66× 16.0× 29%
RX 6900 XT 80 2,885 0.063 19.95× 40.0× 50%

Analysis:

RDNA 2 CU scaling is heavily sub-linear at 1M particles. The RX 6600 XT (16x the CUs) achieves only 4.66x the compute throughput of the iGPU, a 29% scaling efficiency. The RX 6900 XT (40x the CUs) reaches 19.95x, or 50% efficiency -- better than the 6600 XT because its 512 GB/s bandwidth (vs 256 GB/s) helps feed the additional CUs.

The primary bottleneck is memory bandwidth. At 1M particles (32 MB SSBO), the working set is small enough to benefit from cache on the iGPU but must stream through GDDR6 on the discrete cards. The iGPU's DDR5 bandwidth (~83 GB/s) is fully utilised by just 2 CUs, while the RX 6600 XT's 256 GB/s must be shared across 32 CUs. At higher particle counts (16M+), scaling efficiency would improve as the per-CU workload increases and fixed dispatch overhead is amortised.


7d. Architecture Generational Progression

Generational Progression

Normalise all GPUs to the RX 580 (GCN 4) = 1.00× baseline for generational comparison.

Windowed 1M particles (presentation-limited for fast GPUs):

GPU Architecture Year FPS vs RX 580 Memory BW BW vs RX 580
HD 5770 TeraScale 2 2009 188 0.21× 76.8 GB/s 0.30×
FirePro D700 GCN 1.0 2013 555 0.61× 264 GB/s 1.03×
RX 580 GCN 4 2017 912 1.00× 256 GB/s 1.00×
Vega FE GCN 5 2017 1,716 1.88× 483 GB/s 1.89×
RX 6600 XT RDNA 2 2021 1,834 2.01× 256 GB/s 1.00×
RX 6900 XT RDNA 2 2020 4,068 4.46× 512 GB/s 2.00×
RX 9070 XT RDNA 4 2025 1,774 1.95× 640 GB/s 2.50×

Headless 1M particles (pure compute, no presentation overhead):

GPU Architecture Headless FPS (best API) vs RX 580 Windowed → Headless Speedup
RX 6900 XT RDNA 2 (80 CU) 15,950 17.49× 3.9×
RX 9070 XT RDNA 4 (64 CU) 21,354 23.41× 12.0×

Analysis:

The windowed 1M table has an inherent limitation: fast GPUs are presentation-limited, so FPS reflects swapchain throughput rather than GPU compute speed. The RX 9070 XT (1.95×) appears slower than the RX 6900 XT (4.46×) despite having faster per-CU compute — because swapchain throttling caps the 9070 XT more severely (its 64 CUs finish compute faster, spending more time waiting for presentation).

The headless data removes this distortion. With presentation overhead eliminated:

  • The RX 9070 XT is the fastest AMD GPU tested at 21,354 FPS — 1.34× faster than the 80-CU RX 6900 XT (15,950 FPS), despite having only 80% of its CU count. This means RDNA 4's per-CU compute efficiency is ~1.7× higher than RDNA 2.
  • The windowed → headless speedup reveals how much performance is hidden by presentation: the 9070 XT unlocks 12× more throughput in headless mode, while the 6900 XT unlocks only 3.9× — because the 6900 XT's 80 CUs already partially saturate the presentation pipeline in windowed mode.

Bandwidth remains the dominant performance predictor on GCN: Vega FE achieves 1.88× the FPS of the RX 580, almost exactly matching its 1.89× bandwidth ratio. The RX 6600 XT breaks this pattern: identical bandwidth to the RX 580 (256 GB/s, 1.00×) yet 2.01× the FPS — evidence that RDNA 2's architectural improvements (better cache, improved scheduler, higher per-CU throughput) contribute beyond raw bandwidth.

TFLOPS remains a poor predictor: Vega FE has 13.1 TFLOPS (2.1× the RX 580's 6.2) but only 1.88× FPS, while the RX 6600 XT has 10.6 TFLOPS (1.7×) but 2.01× FPS. This workload is bandwidth-bound, not compute-bound.


7e. Cross-API Performance Variation by Architecture

Does the API ranking (DX11 > DX12 > Vulkan > OpenGL) observed on RTX 5090 hold across all AMD architectures, or does it change?

GPU Architecture Fastest API DX11 FPS DX12 FPS Vulkan FPS OpenGL FPS DX11 vs Vulkan Gap
HD 5770 TeraScale 2 OpenGL 107 N/A N/A 188 N/A
RX 580 GCN 4 DX12 755 912 783 42 −3.6%
Vega FE GCN 5 DX12 1,436 1,716 1,370 158 +4.8%
RX 6600 XT RDNA 2 DX12 988 1,834 1,239 180 −20.3%
RX 9070 XT RDNA 4 DX11 1,774 1,609 1,751 253 +1.3%
RX 6900 XT RDNA 2 DX11 4,068 3,518 2,885 229 +41.0%
Zen4/5 iGPU (2 CU) RDNA 2 DX12 271 324 238 229 +13.9%

Analysis:

Unlike the RTX 5090 where DX11 was consistently fastest, AMD GPUs show a more varied API ranking. DX12 is the fastest API on four of seven GPUs (RX 580, Vega FE, RX 6600 XT, iGPU), while DX11 wins on the RX 9070 XT and RX 6900 XT. This suggests AMD's DX12 driver is more competitive than NVIDIA's for simple workloads, while DX11 still benefits the fastest GPUs where CPU overhead matters most.

The DX11 vs Vulkan gap is much smaller on AMD than on NVIDIA. The RX 9070 XT shows only +1.3% for DX11 over Vulkan, and the RX 580 shows just -3.6% (Vulkan slightly faster). The largest gap is the RX 6900 XT at +41.0%, where DX11 dramatically outperforms Vulkan -- likely because at 4,068 FPS, every microsecond of per-frame CPU overhead matters, and AMD's DX11 driver path has lower latency than the Vulkan submission path at this extreme frame rate.

The RX 6600 XT shows a notable -20.3% gap (Vulkan faster than DX11), indicating that AMD's DX11 driver for RDNA 2 mid-range parts has higher overhead than the Vulkan path. This contrasts with the 6900 XT result and suggests driver optimisation varies by SKU.

OpenGL is not competitive on any modern AMD GPU, falling to 42 FPS on the RX 580 (21x slower than DX12). The sole exception is the HD 5770 where OpenGL (188 FPS) outperforms DX11 (107 FPS) -- TeraScale 2 predates compute shaders in DX11, and the OpenGL path may use a more efficient fallback. The FirePro D700 is notable for near-identical performance across all four APIs (516-555 FPS), suggesting its GCN 1.0 architecture is purely GPU-bound regardless of API overhead.


7f. Memory Bandwidth as Performance Predictor

Bandwidth Correlation

Plot FPS against memory bandwidth to test the hypothesis that this benchmark is bandwidth-bound.

GPU Memory BW (GB/s) FPS (best API) FPS / (GB/s)
HD 5770 76.8 188 2.45
FirePro D700 264 555 2.10
RX 580 256 912 3.56
Vega FE (HBM2) 483 1,716 3.55
RX 6600 XT 256 1,834 7.16
RX 6900 XT 512 4,068 7.95
RX 9070 XT 640 1,774 2.77
Zen4/5 iGPU (DDR5) ~83 324 3.90

Analysis:

The FPS/(GB/s) ratio reveals two distinct tiers of bandwidth efficiency. Older architectures (HD 5770, FirePro D700, RX 580, Vega FE) cluster around 2.1-3.6 FPS per GB/s, while RDNA 2 parts (RX 6600 XT, RX 6900 XT) achieve 7.2-8.0 FPS per GB/s -- roughly 2x the bandwidth efficiency. This confirms that RDNA 2's Infinity Cache dramatically amplifies effective bandwidth for workloads with temporal locality, as the 32 MB SSBO working set partially fits in the cache.

Vega FE (3.55) matches RX 580 (3.56) almost exactly in FPS per GB/s despite being a different architecture (GCN 5 vs GCN 4). This means Vega FE's 1.88x FPS advantage comes almost entirely from its 1.89x bandwidth advantage (HBM2), with negligible architectural efficiency gain for this workload.

The RX 9070 XT (3.46) falls to the GCN-era efficiency tier despite being RDNA 4, because its FPS is presentation-limited at 1,774 FPS. The GPU finishes compute in 0.033 ms but waits for swapchain presentation, wasting most of its bandwidth potential.

The iGPU (3.90) performs slightly above the GCN tier despite using shared DDR5. Its 2 CUs generate low enough memory traffic that DDR5 bandwidth is not contended with CPU traffic, and the unified memory architecture avoids PCIe overhead entirely.


7g. Particle Count Scaling — GPU-Bound vs CPU-Bound Crossover

Run each GPU at multiple particle counts to find the crossover point where the bottleneck shifts from CPU to GPU.

GPU Architecture 1M FPS (best API) 16M FPS (best API) 16×/1M Ratio
RTX 5090 Blackwell (170 SM) 7,737 612 12.6×
RX 9070 XT RDNA 4 (64 CU) 1,774 111 16.0×
RX 6900 XT RDNA 2 (80 CU) 4,068 218 18.7×
RX 6600 XT RDNA 2 (32 CU) 1,834 — —
Vega FE GCN 5 (64 CU) 1,716 — —
RX 580 GCN 4 (36 CU) 912 — —
HD 5770 TeraScale 2 188 — —
Zen4/5 iGPU (2 CU) RDNA 2 324 — —

Note: The RTX 5090 ($1,999, flagship tier) is included as a cross-vendor reference point, not as a direct competitor to the mid-range RX 9070 XT ($599) or RX 6900 XT (launched $999, now ~$400 used). Price-performance analysis is in Section 8.

Analysis:

At 1M particles, the fast GPUs are presentation-limited — FPS reflects swapchain throughput rather than GPU compute speed. The 16M column removes this limitation and reveals true GPU-bound performance.

  • RX 6900 XT leads among AMD GPUs at 16M (218 FPS vs 111 FPS for the 9070 XT). This is not a compute advantage — the 9070 XT is faster in pure compute (headless: 21K vs 15K FPS). The difference is the render pass: 16M point primitives stress raw rasterisation throughput, where the 6900 XT's 80 CUs provide 1.25× more parallel rasterisation hardware than the 9070 XT's 64 CUs (detailed analysis in Section 5b).
  • Scaling ratios vary by architecture: the RX 6900 XT's 18.7× ratio (vs expected 16×) reflects its 1M score being more presentation-limited than the 9070 XT's. The RTX 5090's 12.6× suggests its 1M score was already partially GPU-bound (less room to "unlock").
  • RTX 5090 at 16M (612 FPS) is 2.8× the RX 6900 XT — a substantial gap, but narrower than the ~4× implied by their TFLOPS ratio (105 vs 23 TFLOPS), confirming this workload is bandwidth-bound rather than compute-bound.

7h. RenderDoc Cross-GPU Frame Analysis

Capture one Vulkan frame on each AMD GPU (where Vulkan is available) and compare per-event GPU timing, barrier cost, and command structure.

Captures generated via --capture 5 (auto-capture at 5 seconds). Analysis automated with scripts/rdoc_analyse.py and scripts/rdoc_export_timing.py.

Per-Event GPU Timing Comparison

GPU Architecture Compute Dispatch (ms) Barrier (ms) Render Pass (ms) Total GPU (ms)
RX 9070 XT RDNA 4 (64 CU) 0.033 0.005 0.408 0.446
RX 6900 XT RDNA 2 (80 CU) 0.063 < 0.001 0.133 0.196
RX 6600 XT RDNA 2 (32 CU) 0.270 < 0.001 0.379 0.649
Vega FE GCN 5 (64 CU) 0.219 0.001 0.275 0.495
RX 580 GCN 4 (36 CU) 0.362 0.006 0.702 1.070
FirePro D700 GCN 1.0 (12 CU) 0.589 0.003 0.881 1.473
Zen4/5 iGPU (2 CU) RDNA 2 1.257 < 0.001 1.928 3.185

HD 5770 excluded — no Vulkan support. Barrier cost derived from Total GPU − Compute − Render (timestamps 1→4 minus timestamps 1→2 and 3→4). Values < 0.001 ms are below timestamp query resolution.

App Timestamp vs RenderDoc Cross-Validation

GPU Metric App Timestamp (ms) RenderDoc CPU Trace (µs) Notes
RX 9070 XT Compute 0.033 0.009 CPU recording ≪ GPU execution
RX 9070 XT Render 0.408 0.051 GPU-bound render pass
RX 6900 XT Compute 0.063 0.004 CPU recording ≪ GPU execution
RX 6900 XT Render 0.133 0.074 GPU-bound render pass
RX 6600 XT Compute 0.270 0.004 CPU recording ≪ GPU execution
RX 6600 XT Render 0.379 0.076 GPU-bound render pass
Vega FE Compute 0.219 0.005 CPU recording ≪ GPU execution
Vega FE Render 0.275 0.099 GPU-bound render pass
RX 580 Compute 0.362 0.004 CPU recording ≪ GPU execution
RX 580 Render 0.702 0.075 GPU-bound render pass
Zen4/5 iGPU (2 CU) Compute 1.257 0.005 CPU recording ≪ GPU execution
Zen4/5 iGPU (2 CU) Render 1.928 0.107 GPU-bound render pass

Note: RenderDoc JSON exports record CPU-side API call durations (when each vkCmd* was recorded into the command buffer), not GPU execution time. GPU-side per-event timing requires opening the .rdc files in the RenderDoc GUI Performance Counter Viewer. The CPU trace confirms all GPUs are fully GPU-bound: CPU recording is 100–1000× faster than GPU execution for every event.

Barrier Cost Comparison

Barrier Cost

GPU Architecture Memory Type Barrier Duration (ms) Notes
RX 9070 XT RDNA 4 16 GB GDDR6 0.005 Measurable — RDNA 4 L2 writeback
RX 6900 XT RDNA 2 16 GB GDDR6 < 0.001 Below timestamp resolution
RX 6600 XT RDNA 2 8 GB GDDR6 < 0.001 Below timestamp resolution
Vega FE GCN 5 16 GB HBM2 0.001 HBM2 — near-zero
RX 580 GCN 4 8 GB GDDR5 0.006 Highest — GDDR5 L2 flush
FirePro D700 GCN 1.0 6 GB GDDR5 0.003 Moderate — early GCN
Zen4/5 iGPU (2 CU) RDNA 2 Shared DDR5 < 0.001 Unified memory — near-zero

Analysis:

Barrier cost is negligible across all tested AMD GPUs (< 0.006 ms), confirming that the VK_PIPELINE_STAGE_COMPUTE_SHADER_BIT → VK_PIPELINE_STAGE_VERTEX_INPUT_BIT buffer memory barrier introduces minimal synchronisation overhead. The RX 580 (GCN 4, GDDR5) shows the highest measurable barrier at 0.006 ms, likely due to its older L2 cache architecture requiring a full writeback. RDNA 2 GPUs (6900 XT, 6600 XT, iGPU) all show sub-microsecond barriers, suggesting the Infinity Cache absorbs the coherency cost. The iGPU's unified memory architecture confirms the expected near-zero barrier cost — no physical cache flush is needed when compute and graphics share the same memory controller.

Per-Frame Event Count by GPU

GPU Total Events Frame Events Dispatches Draw Calls Barriers Driver-Inserted
RX 9070 XT ~135 52 1 1 1 0
RX 6900 XT ~130 50 1 1 1 0
RX 6600 XT ~135 51 1 1 1 0
Vega FE ~130 52 1 1 1 0
RX 580 ~130 50 1 1 1 0
FirePro D700 ~130 50 1 1 1 0
Zen4/5 iGPU (2 CU) ~135 52 1 1 1 0

All AMD GPUs produce an identical frame structure: 1 compute dispatch, 1 pipeline barrier, 1 render pass (begin + draw + end), 4 timestamp writes, and 3 debug label pairs. Frame event counts (50–52) vary slightly due to some debug label events being recorded as instant events vs begin/end pairs across driver versions. No driver-inserted implicit barriers observed on any AMD GPU — the single explicit vkCmdPipelineBarrier is sufficient.

NVIDIA RTX 5090 comparison (Section 3): 133 total events, 28 frame events — the higher total reflects NVIDIA's driver inserting additional internal resource tracking events outside the frame boundary.


7i. 3DMark Cross-Validation — AMD GPU Fleet

Cross-validate this benchmark's AMD GPU rankings against 3DMark Time Spy (DX12) and Fire Strike (DX11) to confirm the results reflect real-world performance scaling.

Data source: scripts/3dmark_scores.json Charts generated with: python scripts/compare_3dmark.py --save docs/images

Normalised Performance (Baseline: RX 580 = 1.00×)

GPU Architecture This Benchmark 3DMark Time Spy 3DMark Fire Strike Deviation (TS) Deviation (FS)
RX 9070 XT RDNA 4 (64 CU) 1.95× 6.49× 4.44× −70.0% −56.1%
RX 6900 XT RDNA 2 (80 CU) 4.46× 4.63× 3.82× −3.7% +16.8%
RX 6600 XT RDNA 2 (32 CU) 2.01× 2.16× 1.92× −7.0% +4.9%
Vega FE GCN 5 (64 CU) 1.88× 1.59× 1.46× +18.2% +28.8%
RX 580 GCN 4 (36 CU) 1.00× 1.00× 1.00× — —
Zen4/5 iGPU (2 CU) RDNA 2 0.36× 0.16× 0.15× +119.2% +132.0%
HD 5770 TeraScale 2 0.21× N/A 0.10× N/A +106.8%

Deviation = (This Benchmark ratio / 3DMark ratio) − 1. Positive = our benchmark favours that GPU more; negative = less.

Expected Deviations and Why

GPU Expected Deviation Reason
RX 9070 XT Negative (−50–70%) FPS is presentation-limited at 1,774 FPS; 3DMark exercises full GPU feature set
RX 6900 XT Near zero (TS) / Positive (FS) 80 CU + 512 GB/s bandwidth scales well for both workloads; Fire Strike's DX11 overhead less efficient than Vulkan compute
RX 6600 XT Near zero Mid-range GPU; balanced for both workload types
Vega FE Positive (+18–29%) HBM2's 483 GB/s bandwidth disproportionately benefits bandwidth-bound compute workloads
Zen4/5 iGPU (2 CU) Positive (+119–132%) Simple compute fits within 2-CU cache; 3DMark's complex workloads expose shader count limit
HD 5770 N/A for Time Spy TeraScale 2 has no DX12; Fire Strike deviation +107% (similar to iGPU — simple compute overperforms)

Correlation Analysis

The deviations reveal how this benchmark's single-dispatch compute workload differs from 3DMark's complex multi-pass rasterisation:

  • RX 9070 XT (−70.0% TS, −56.1% FS): The largest deviation. In 3DMark, the 9070 XT exercises its full RDNA 4 feature set (mesh shaders, ray tracing units, 64 CUs at full utilisation). In this benchmark, windowed FPS is presentation-limited at 1,774 FPS — the GPU finishes compute in 0.033 ms but waits for swapchain presentation (Section 6). Headless mode resolves this: at 21,354 FPS (23.41× vs RX 580), the deviation flips to +261% (TS) and +427% (FS) — confirming that when presentation overhead is removed, this compute benchmark massively favours the 9070 XT over 3DMark's rasterisation workloads. The windowed deviation measures presentation bottleneck; the headless deviation measures the fundamental compute-vs-rasterisation gap.
  • Vega FE (+18.2% TS, +28.8% FS): HBM2's 483 GB/s bandwidth disproportionately benefits this bandwidth-bound compute workload. 3DMark's texture-heavy scenes do not benefit as much from raw bandwidth.
  • iGPU (+119% TS, +132% FS): The 2-CU iGPU overperforms relative to 3DMark because this benchmark's simple compute workload fits well within the iGPU's cache and DDR5 bandwidth is adequate for 2 CUs. 3DMark's complex geometry and texture workloads expose the iGPU's limited shader count.
  • RX 6900 XT (−3.7% TS) and RX 6600 XT (−7.0% TS): Within the expected ±10% range for Time Spy in windowed mode, confirming that RDNA 2 GPUs are well-correlated. In headless mode, the 6900 XT jumps to 17.49× (vs 4.46× windowed), yielding +278% (TS) and +358% (FS) deviations — similar to the 9070 XT pattern, confirming that all fast GPUs are presentation-limited in windowed mode. Fire Strike deviations are slightly higher (+16.8% and +4.9%) because DX11's heavier CPU overhead in 3DMark penalises complex scenes more than this benchmark's single-dispatch Vulkan path.
  • HD 5770 (+106.8% FS): Similar to the iGPU — TeraScale 2's limited shader hardware handles this simple compute workload relatively better than 3DMark's complex rasterisation. No Time Spy comparison possible (no DX12).

The large deviations for the 9070 XT (presentation-limited) and the low-end GPUs (iGPU, HD 5770) confirm that this benchmark measures a fundamentally different aspect of GPU performance (bandwidth-bound compute + presentation overhead) compared to 3DMark (shader-heavy rasterisation). GPUs in their GPU-bound regime (6900 XT, 6600 XT) show strong correlation. Both benchmarks are needed for a complete performance picture.

This Benchmark vs 3DMark — What They Measure Differently

This Benchmark 3DMark Time Spy 3DMark Fire Strike
Workload Single compute dispatch + single draw call Multi-pass rasterisation, tessellation, post-FX Multi-pass rasterisation, particle physics
Draw calls / frame 1 Thousands Thousands
Bottleneck (fast GPU) CPU overhead GPU (texture, geometry, shading) GPU (texture, shading)
Memory access Sequential SSBO read/write Random texture fetches, render targets Random texture fetches
Benefits from Memory bandwidth, compute scheduler Shader count, TMUs, ROPs, driver DX12 path Shader count, TMUs, ROPs, driver DX11 path

This difference explains why deviations exist — and why both benchmarks are needed for a complete GPU performance picture.


7j. 3DMark API Overhead — Draw Call Throughput

The 3DMark API Overhead test measures raw draw call throughput (draw calls per second) for each graphics API. This directly complements this benchmark's cross-API performance comparison by isolating driver/API overhead from GPU compute capability.

Data source: 3DMark Results/*/...-api-result.3dmark-result → extracted via scripts/extract_3dmark.py

GPU Specifications and Draw Call Throughput

GPU Architecture FP32 (TFLOPS) Boost Clock DX11-ST DX11-MT DX12 Vulkan VK / DX11
RX 9070 XT RDNA 4 (64 CU) 48.7 † 2970 MHz 2.41 M 3.05 M 25.88 M 39.88 M 16.5×
RX 6900 XT RDNA 2 (80 CU) 23.04 2250 MHz 2.48 M 3.12 M 40.78 M 36.75 M 14.8×
RX 6600 XT RDNA 2 (32 CU) 10.60 2589 MHz 2.51 M 3.19 M 40.32 M 39.25 M 15.6×
Vega FE GCN 5 (64 CU) 13.11 1600 MHz 2.22 M 2.49 M 26.57 M 26.06 M 11.8×
RX 580 GCN 4 (36 CU) 6.17 1340 MHz 2.38 M 2.30 M 27.67 M 26.92 M 11.3×
FirePro D700 GCN 1.0 (32 CU) 3.48 850 MHz 1.16 M 1.03 M 10.04 M 9.46 M 8.2×
HD 5770 TeraScale 2 (800 SP) 1.36 850 MHz 1.60 M 1.60 M — — —
Ryzen 5 7600 iGPU RDNA 2 (2 CU) 0.56 2200 MHz 2.43 M 3.05 M 7.22 M 7.14 M 2.9×
Adreno 640 Adreno 6xx (768 ALU) ~0.90 585 MHz 0.16 M 0.18 M 0.87 M 0.70 M 4.3×

† RDNA 4 uses dual-issue FP32 (each stream processor executes 2 FP32 ops/clock). Without dual-issue, the traditional calculation yields 24.3 TFLOPS — comparable methodology to the RDNA 2 and GCN numbers above.

Draw call throughput is in millions of draw calls per second. DX11-ST = single-threaded, DX11-MT = multi-threaded.

Sources: AMD official specifications, TechPowerUp GPU Database, 3DMark API Overhead test results.

Key Observations

  1. DX11 is CPU-bound, not GPU-bound: All discrete GPUs from HD 5770 to RX 9070 XT cluster at 1.6–2.5 M DX11 single-threaded draw calls regardless of GPU power. The bottleneck is the single-threaded CPU driver path, not the GPU. DX11 multi-threading provides only 1.0–1.3× improvement on AMD (compared to NVIDIA's often 0.1× due to different driver threading models).

  2. RDNA 2 has the best DX12 driver efficiency: The RX 6900 XT (40.78 M) and RX 6600 XT (40.32 M) both outperform the RX 9070 XT (25.88 M) in DX12 draw call throughput. This is likely because the RDNA 4 DX12 driver is newer and less optimised — the same pattern seen in early driver releases for previous AMD architectures.

  3. RDNA 4 leads in Vulkan: The RX 9070 XT (39.88 M) achieves the highest Vulkan throughput, slightly above the RX 6600 XT (39.25 M). This aligns with this benchmark's Vulkan results where the 9070 XT performs best.

  4. GCN 4/5 parity: The RX 580 (GCN 4) and Vega FE (GCN 5) show nearly identical API overhead (~27 M DX12, ~26 M Vulkan), confirming that the driver stack is the same between these two GCN generations.

  5. iGPU bandwidth-limited: The Ryzen 5 7600 iGPU achieves only 7 M DX12/Vulkan despite using the same RDNA 2 driver as the 6600 XT (40 M). The 2-CU iGPU's shared DDR5 memory bandwidth (not driver overhead) limits draw call submission rate.

  6. GCN 1.0 (D700) shows age: At 10 M DX12 and 9.5 M Vulkan, the 2013-era FirePro D700 achieves 25% of RDNA 2's throughput — a reasonable result given the 8-year architecture gap and older driver codepath.

  7. TeraScale 2 is DX11-only: The HD 5770 has no DX12 or Vulkan support, so API overhead comparison is limited to DX11 where it achieves 1.60 M (67% of modern GPUs' DX11 rate — bottlenecked by the older CPU driver).

  8. Adreno 640 (mobile SoC): At 0.16 M DX11-ST and 0.70 M Vulkan, the Snapdragon 855's GPU achieves only 7% of desktop AMD DX11 throughput and 2% of desktop Vulkan throughput. The 4.3× VK/DX11 ratio is much lower than desktop GPUs (11–17×), reflecting the mobile driver's relatively efficient DX11 path (via translation layer) and limited Vulkan command processor bandwidth.

Correlation with This Benchmark's Cross-API Results

The API Overhead results explain patterns observed in this benchmark's cross-API testing (Section 5):

  • Why Vulkan ≈ DX11 on AMD 9070 XT: The 9070 XT's Vulkan driver (39.88 M) is 16.5× faster than DX11 (2.41 M) at draw call submission. But this benchmark's single draw call per frame means the per-call overhead difference is negligible — both APIs spend < 0.01 ms on the draw call itself. The performance similarity is expected for single-dispatch workloads.

  • Why DX12 underperforms on 9070 XT: Despite DX12's explicit nature, the 9070 XT's DX12 driver (25.88 M) achieves only 65% of its Vulkan throughput. This immature DX12 driver may also explain the slightly lower DX12 FPS observed in this benchmark's cross-API comparison.


7k. Summary — AMD Generational Findings

Observation Explanation
RDNA 4 (9070 XT) achieves 4.1× better per-CU compute than RDNA 2 (6600 XT) with 2× the CU count (64 vs 32) Higher clocks (1.15×) + architectural improvements in scheduler, cache hierarchy, and driver codegen account for the remaining ~1.8×
9070 XT outperforms 80-CU RX 6900 XT in compute (0.033 vs 0.063 ms) despite 80% CU count RDNA 4's per-CU efficiency (~1.7× higher) is enough to overcome the 1.25× CU disadvantage
API ranking on 9070 XT (DX11 ≈ Vulkan > DX12 > OpenGL) differs from RTX 5090 (DX11 >> DX12 > Vulkan > OpenGL) AMD's DX11 driver is less optimised than NVIDIA's; the gap between explicit and implicit APIs is much smaller on AMD
OpenGL compute overhead persists on RDNA 4 (~2.6 ms) at similar levels to RDNA 2 (~2.7 ms) AMD's OpenGL-to-Vulkan translation layer has not improved compute dispatch overhead across generations
9070 XT compute scales 59× for 16× particle increase (1M → 16M) Super-linear: at 1M particles GPU is underutilised, per-dispatch overhead dominates; at 16M the ALUs and bandwidth are fully saturated
TeraScale 2 VLIW5 achieves only ~50–70% of theoretical per-SP throughput VLIW5 slot packing inefficiency in compute shaders with irregular control flow
All APIs converge to 0.034 ms compute in headless mode on 9070 XT Proves GPU compute hardware is identical across APIs; windowed differences are entirely presentation/driver overhead

8. Cross-Validation Against 3DMark

To confirm that this benchmark accurately reflects real-world GPU performance differences, results are cross-validated against 3DMark — the industry-standard graphics benchmark by UL (formerly Futuremark).

Methodology

  1. Run this project's benchmark on each GPU (best FPS across all APIs).
  2. Run 3DMark Time Spy (DX12) and Fire Strike (DX11) on the same GPUs.
  3. Normalise all scores to a common baseline GPU (e.g. RX 580 = 1.00×).
  4. Compare the relative performance ratios.

If both benchmarks rank GPUs in the same order with similar ratios, it validates that this project's compute-heavy workload is a meaningful GPU performance indicator.

Normalised Performance Comparison

Baseline: RX 580 = 1.00×

GPU Architecture This Benchmark 3DMark Time Spy 3DMark Fire Strike Deviation (TS) Deviation (FS)
RTX 5090 Blackwell (170 SM) 8.49× 8.66× 5.62× −2.0% +51.1%
RX 9070 XT RDNA 4 (64 CU) 1.95× 6.49× 4.44× −70.0% −56.1%
RX 6900 XT RDNA 2 (80 CU) 4.46× 4.63× 3.82× −3.7% +16.8%
RX 6600 XT RDNA 2 (32 CU) 2.01× 2.16× 1.92× −7.0% +4.9%
Vega FE GCN 5 (64 CU) 1.88× 1.59× 1.46× +18.2% +28.8%
RX 580 GCN 4 (36 CU) 1.00× 1.00× 1.00× — —
Zen4/5 iGPU (2 CU) RDNA 2 0.36× 0.16× 0.15× +119.2% +132.0%
HD 5770 TeraScale 2 0.21× N/A (no DX12) 0.10× N/A +106.8%

Deviation = (This Benchmark ratio / 3DMark ratio) − 1. Positive means our benchmark favours that GPU more than 3DMark; negative means less.

Expected Deviations and Why

This project runs a single compute dispatch + single draw call per frame. 3DMark runs complex multi-pass rasterisation with thousands of draw calls, tessellation, post-processing, and full-screen effects. Expected differences:

GPU Deviation (TS) Deviation (FS) Reason
RTX 5090 −2.0% +51.1% Time Spy near-perfect; Fire Strike deviation due to DX11 CPU overhead in 3DMark vs single-dispatch Vulkan here
RX 9070 XT −70.0% −56.1% FPS is presentation-limited (swapchain throttling) at 1,774 FPS despite 0.033 ms compute. 3DMark exercises the full GPU; this benchmark cannot (see Section 6)
RX 6900 XT −3.7% +16.8% Strong Time Spy correlation; Fire Strike deviation from DX11 overhead differential
RX 6600 XT −7.0% +4.9% Good correlation across both benchmarks
Vega FE +18.2% +28.8% HBM2's 483 GB/s bandwidth disproportionately benefits this bandwidth-bound compute workload; 3DMark's texture-heavy scenes do not leverage raw bandwidth as heavily
Zen4/5 iGPU (2 CU) +119.2% +132.0% The iGPU's 2 CUs handle this simple compute workload efficiently (cache-friendly, low contention), but 3DMark's complex geometry/texture workloads expose the severe shader count limitation
HD 5770 N/A +106.8% TeraScale 2 has no DX12; Fire Strike deviation similar to iGPU — simple compute overperforms relative to 3DMark's complex rasterisation

GPUs operating in their GPU-bound regime (RTX 5090, RX 6900 XT, RX 6600 XT) show Time Spy deviations within ±10%, confirming strong correlation. The large deviations for the RX 9070 XT (presentation-limited), Vega FE (bandwidth advantage), iGPU and HD 5770 (workload mismatch) are explainable by workload characteristics and are not indicative of benchmark error.

Correlation Analysis

A linear regression of project FPS vs 3DMark Time Spy scores across all GPUs yields a strong linear relationship for most GPUs, with the RX 9070 XT and iGPU as known outliers due to presentation throttling and workload mismatch respectively (see Section 6). Excluding these outliers, the remaining GPUs (RTX 5090, RX 6900 XT, RX 6600 XT, RX 580) show Time Spy deviations within ±10%, confirming the benchmark's validity for GPUs operating in their GPU-bound regime.

Charts: Run python scripts/compare_3dmark.py --save docs/images to generate the normalised bar chart and correlation scatter plot (docs/images/3dmark_comparison.png, docs/images/3dmark_correlation.png).

Data Source

3DMark scores are stored in scripts/3dmark_scores.json. To auto-import from 3DMark result files:

# Import from .3dmark-result files (3DMark Advanced/Professional)
python scripts/compare_3dmark.py --import-3dmark "C:\Users\*\Documents\3DMark\*.3dmark-result"

The .3dmark-result format is a ZIP archive containing arielle.xml (benchmark scores, per-loop FPS) and si.xml (GPU name, VRAM, driver version). The import script parses both and merges into the JSON scores file.


System Configuration

Component Specification
CPU AMD Ryzen 7 9800X3D 8-Core Processor
Discrete GPU NVIDIA GeForce RTX 5090 (32 GB GDDR7)
Integrated GPU AMD Radeon Graphics (Zen 4, 2 CU, 2 GB shared DDR5)
OS Windows 11 25H2
Resolution 1280 × 720
V-Sync OFF
Memory Mode Device-local (staging buffer → VRAM on dGPU)

Part 4 — Other GPUs and Cross-Platform

9. Cross-API Comparison — RTX 5090, 1M Particles (Medium)

Metric Vulkan DirectX 12 DirectX 11 OpenGL 4.3
Avg FPS 3,611 6,547 8,955 2,442
Avg GPU Time 0.094 ms 0.065 ms 0.104 ms 0.087 ms
Avg Frame Time 0.277 ms 0.153 ms 0.112 ms 0.409 ms
CPU Overhead / Frame 0.183 ms 0.088 ms 0.008 ms 0.322 ms
GPU Utilisation 33.9% 42.4% 93.2% 21.2%
Bottleneck CPU-bound CPU-bound GPU-bound CPU-bound

Key Finding: DX11 Achieves Highest FPS for Simple Workloads

All four APIs deliver nearly identical GPU execution times (0.065–0.104 ms), confirming the GPU-side workload is equivalent. The FPS difference is entirely driven by per-frame CPU overhead.

Ranking: DX11 > DX12 > Vulkan > OpenGL

Why DX11 is fastest here:

  • DX11 is an implicit API — the driver handles command batching, resource state tracking, and barrier insertion internally. NVIDIA's DX11 driver path has been optimised for over a decade, making it extremely efficient for simple, single-threaded workloads.
  • Per-frame CPU overhead is only 0.008 ms, leaving the GPU as the actual bottleneck (93.2% utilisation).

Why DX12/Vulkan are slower here:

  • Both are explicit APIs requiring the application to manually manage command allocators, fences, resource barriers, and descriptor heaps/sets.
  • This shifts work from the driver to application code, adding 10–20× more CPU overhead per frame compared to DX11.
  • With only 1 compute dispatch + 1 draw call per frame, there is no opportunity for multi-threaded command recording — the very feature that justifies explicit APIs in complex scenes.

Why OpenGL is slowest here:

  • OpenGL has the highest per-frame CPU overhead at 0.322 ms — roughly 40× more than DX11.
  • glfwSwapBuffers on Windows goes through WGL, which has less efficient frame queue management than DXGI's Present path.
  • OpenGL's global state machine model means every glUseProgram, glBindBuffer, and glBindVertexArray call triggers internal driver state validation, accumulating significant overhead even with minimal draw calls.
  • Despite being an implicit API like DX11, OpenGL's Windows driver path has received far less optimisation from NVIDIA in recent years, as industry focus has shifted to Vulkan and DirectX.
  • However, OpenGL still produces the second-fastest GPU execution time (0.087 ms), confirming the bottleneck is purely in the CPU-side driver, not in the shader or buffer management.

When DX12/Vulkan win:

Explicit APIs excel when a scene contains hundreds or thousands of draw calls. In that scenario, DX11's single-threaded driver becomes the bottleneck, while DX12/Vulkan can parallelise command recording across multiple CPU threads, reducing total CPU time proportionally.

When OpenGL makes sense:

OpenGL 4.3 remains the most portable option — it runs on Windows, Linux, and macOS (legacy profile) without requiring Vulkan drivers or platform-specific APIs. For workloads that are GPU-bound (high particle counts, complex shaders), OpenGL's higher CPU overhead becomes negligible relative to total frame time.

AMD comparison: The RX 9070 XT (RDNA 4) shows a very different API ranking: DX11 ≈ Vulkan (1,774 / 1,751 FPS) > DX12 (1,609 FPS) >> OpenGL (253 FPS). The DX11-over-Vulkan advantage shrinks from 2.5× on NVIDIA to 1.01× on AMD, reflecting AMD's less optimised DX11 driver. See Section 5 for the full RX 9070 XT cross-API analysis.


10. Cross-GPU Comparison — Vulkan, 1M Particles (Medium), Device-local

Metric RTX 5090 (Discrete) RX 9070 XT (Discrete) AMD Radeon iGPU (Integrated)
Avg FPS 2,700+ 1,751 ~320
Compute 0.035 ms 0.033 ms 1.47 ms
Render 0.045 ms 0.408 ms 1.5 ms
Total GPU 0.08 ms 0.446 ms ~3.0 ms
Ratio 1× ~5.6× slower ~37× slower

The RTX 5090 (21,760 CUDA cores, ~3,000 GB/s bandwidth) outperforms the Zen 4 iGPU (128 shaders, ~50 GB/s shared DDR5) by approximately 37× in GPU execution time. This aligns with the memory bandwidth ratio (~60×), confirming the particle simulation is bandwidth-bound rather than compute-bound at this scale.

The RX 9070 XT sits between these extremes: its compute time (0.033 ms) is comparable to the RTX 5090 (0.035 ms), but its total GPU time (0.446 ms) is 5.6× higher due to swapchain semaphore wait pollution inflating the render timestamp (see Section 6). In headless mode, the 9070 XT achieves 21,260 FPS — only ~15% behind the RTX 5090's headless throughput.

RTX 5090 — 16M Particles (Ultra), Windowed

# API Avg FPS Compute (ms) Render (ms) Total GPU (ms)
1 DX12 611.7 0.581 0.730 1.312
2 Vulkan 539.9 0.628 1.006 1.635
3 DX11 470.0 0.605 0.771 2.000
4 OpenGL 457.6 0.580 0.969 1.550

At 16M particles the RTX 5090 remains fast enough that DX12 retakes the lead from DX11 (612 vs 470 FPS). The workload is now GPU-bound, so DX11's low CPU overhead advantage disappears and its higher total GPU time (2.000 ms, likely due to implicit barrier overhead) becomes the bottleneck.

RTX 5090 — Headless Compute, 1M Particles

# API Avg FPS Compute (ms) Total GPU (ms)
1 DX11 37,564 0.000 0.000
2 OpenGL 35,110 0.018 0.020
3 Vulkan 26,358 0.023 0.025
4 DX12 24,758 0.014 0.014

DX11 timestamp anomaly: The RTX 5090's DX11 headless mode reports 0.000 ms for all GPU timing metrics, suggesting NVIDIA's DX11 driver does not support timestamp queries in headless/compute-only mode. The 37,564 FPS figure is valid (derived from CPU-side frame timing), but the GPU time breakdown is unavailable.

Cross-vendor headless comparison (1M particles, best API):

GPU Best Headless FPS Best API Price (MSRP) FPS per $
RTX 5090 37,564 DX11 $1,999 18.8
RX 9070 XT 21,354 DX12 $599 35.6
RX 6900 XT 15,950 DX12 $999 (launched) 16.0

The RTX 5090 achieves 1.76× the headless throughput of the RX 9070 XT — a significant lead, but far from the ~2.2× TFLOPS ratio (105 vs 48.7 TFLOPS) or the ~2.8× bandwidth ratio (1,792 vs 640 GB/s). On a price-performance basis, the RX 9070 XT delivers 1.9× the FPS per dollar of the RTX 5090 for this compute workload. Even comparing Vulkan-to-Vulkan (26,358 vs 21,260 FPS), the 5090's lead narrows to just 1.24×.

This confirms that for bandwidth-bound compute workloads, mid-range GPUs offer substantially better value than flagships — the RTX 5090's additional CUDA cores and memory bandwidth face diminishing returns when the workload cannot saturate them.


11. Memory Allocation Impact — Vulkan, RTX 5090

Memory Mode Compute Render Total GPU FPS
Device-local (default) 0.035 ms 0.045 ms 0.08 ms 2,700+
Host-visible (--host-memory) 1.25 ms 0.15 ms 1.4 ms ~600

Using host-visible memory (system RAM accessed over PCIe) instead of device-local VRAM causes a 35× increase in compute time on a discrete GPU. The compute shader reads/writes particle data every frame — over PCIe, this becomes the dominant bottleneck.

On an integrated GPU, this penalty disappears because host-visible and device-local memory both reside in the same physical DDR5, making the distinction meaningless.


12. Software Renderer Baseline — WARP, 1M Particles

WARP (Windows Advanced Rasterisation Platform) is Microsoft's CPU-based software rasteriser bundled with every modern Windows installation. It runs the entire graphics pipeline on the CPU using SIMD (SSE/AVX) and multi-threading, serving as both a correctness reference and a fallback when no hardware GPU driver is available.

Native API support: WARP natively implements Direct3D 11 and Direct3D 12 only. It does not implement Vulkan or OpenGL. If Vulkan is reported as available on a WARP-only system, this is provided by Mesa Dozen — a Vulkan-on-D3D12 translation layer distributed via the Microsoft Store's OpenCL, OpenGL & Vulkan Compatibility Pack. Similarly, OpenGL support on WARP comes from OpenGLOn12 in the same compatibility pack. Both layers translate their respective API calls to D3D12, which WARP then executes on the CPU.

12a. Hardware GPU vs WARP

Metric RTX 5090 / DX12 (Hardware) WARP / DX12 (Software)
Avg FPS 6,547 83
Compute 0.014 ms 1.1 ms
Render 0.050 ms 10.6 ms
Total 0.065 ms 11.7 ms

WARP demonstrates a ~80× performance gap compared to hardware GPU execution, which is expected for CPU-based software rasterisation.

12b. WARP: DX11 vs DX12

Metric WARP + DX11 WARP + DX12
Avg FPS 52 83
Timestamp queries Not available 11.7 ms total

On hardware GPUs, DX11 outperforms DX12/Vulkan because the driver's implicit state management is highly optimised and adds negligible overhead. On WARP, the result reverses: DX12 is 60% faster than DX11.

Why the reversal:

  • DX11's driver layer becomes pure overhead. On a hardware GPU, the DX11 runtime performs implicit resource state tracking, dependency analysis, and barrier insertion to optimise GPU command submission. When the "GPU" is WARP (a CPU-based software renderer), there is no hardware to optimise for — this entire layer is wasted CPU work.
  • DX12's thin runtime is a better fit. DX12's explicit model has minimal runtime between the application and the execution engine. The application specifies exactly what to do, and WARP executes it directly with less translation overhead.
  • WARP's DX12 implementation is more modern. DX12 (introduced 2015) benefits from a newer WARP backend that may leverage more efficient internal scheduling compared to the legacy DX11 WARP path.

This observation reinforces that the DX11 driver's "free optimisation" is specifically valuable for hardware GPU command submission — when that hardware is absent, the optimisation layer becomes a liability.

12c. WARP as a Vulkan Device — Dozen Translation Layer

On systems without a native Vulkan ICD (e.g. Windows on ARM VMs, virtual GPUs), the Vulkan loader may enumerate "Microsoft Basic Render Driver" as a Vulkan physical device. The full call chain is:

Vulkan application  →  Dozen (Vulkan → D3D12)  →  WARP (D3D12 → CPU)

Dozen is distributed as part of the OpenCL, OpenGL & Vulkan Compatibility Pack from the Microsoft Store (D3DMappingLayers app package). Windows 11 may install this pack automatically on devices that lack native Vulkan/OpenGL drivers, particularly Windows on ARM devices and virtual machines.

On systems with a hardware Vulkan ICD (e.g. NVIDIA, AMD), the Dozen/WARP device is typically not enumerated or is deprioritised by the Vulkan loader. Selecting Vulkan on a WARP-only system is functionally identical to selecting DX12 on WARP — both end up as CPU-based software rendering, with Dozen adding a thin additional translation layer.


13. Legacy Discrete GPU vs Modern iGPU — Compute Efficiency Beyond TFLOPS

Test System B

Component Specification
CPU AMD Ryzen 5 7600 6-Core Processor
Discrete GPU AMD Radeon HD 5770 (757 MB GDDR5, TeraScale 2, 2009)
Integrated GPU AMD Radeon Graphics (Zen 4 / RDNA 2, 2 CU, shared DDR5)
OS Windows 11 (NT 10.0.26200)
Resolution 1280 × 720
V-Sync OFF
Memory Mode Device-local

The HD 5770 only supports DX11 and OpenGL 4.3 — no Vulkan or DX12 drivers exist for TeraScale 2 hardware. The Ryzen 7600 iGPU supports all four APIs.

13a. Raw Results — 1M Particles (Medium)

# API GPU Avg FPS Compute (ms) Render (ms) Total GPU (ms) Utilisation
1 DX12 Radeon iGPU (RDNA 2) 313 — — 2.956 —
2 Vulkan Radeon iGPU (RDNA 2) 275 — — 3.426 —
3 OpenGL Radeon iGPU (RDNA 2) 275 — — 3.391 —
4 OpenGL HD 5770 (TeraScale 2) 193 1.789 3.025 4.820 93.1%
5 DX11 Radeon iGPU (RDNA 2) 190 — — 5.017 —
6 DX11 HD 5770 (TeraScale 2) 111 1.078 2.715 8.973 99.6%
7 DX12 WARP (CPU) 64 — — 15.234 —
8 DX11 WARP (CPU) 45 — — 21.954 —

13b. Same-API Head-to-Head

API HD 5770 FPS iGPU FPS iGPU Advantage
DX11 111 190 +71%
OpenGL 193 275 +42%

The RDNA 2 integrated GPU outperforms the HD 5770 discrete GPU in every comparable API, despite having far fewer hardware resources on paper.

13c. Theoretical TFLOPS Comparison

HD 5770 Ryzen 5 7600 iGPU
Architecture TeraScale 2 (2009) RDNA 2 (2022)
Stream Processors 800 (160 × VLIW5) 128 (2 CU × 64)
Core Clock 850 MHz 2,200 MHz
FP32 TFLOPS ~1.36 ~0.56
Memory 1 GB GDDR5, ~76.8 GB/s Shared DDR5, ~83 GB/s

The HD 5770 has 2.4× more raw FP32 TFLOPS than the iGPU, yet it is 42–71% slower in this compute benchmark. This inversion demonstrates that TFLOPS alone is a poor predictor of real-world compute shader performance.

This is not anomalous — TFLOPS comparisons across different architectures are unreliable as an industry rule of thumb. Well-documented examples include AMD Vega 64 (13.7 TFLOPS) losing to NVIDIA GTX 1080 (8.9 TFLOPS) in many gaming and compute workloads, and Intel Arc A770 (19.7 TFLOPS) underperforming against the RTX 3060 (12.7 TFLOPS) at launch despite a 55% TFLOPS advantage. TFLOPS measures only the theoretical rate of fused multiply-add operations — it says nothing about whether the ALUs can actually be kept fed with data and useful instructions. Cache hit rates, memory bandwidth, scheduling efficiency, VLIW slot utilisation, and driver code generation quality all determine how much of the theoretical peak is realised in practice.

The discrepancy between TFLOPS rankings and benchmark results is itself a validation of the test. If results tracked TFLOPS perfectly, it would suggest the benchmark is merely saturating ALU throughput with a trivially parallel workload — essentially an artificial peak-FLOPS test. The fact that a 0.56 TFLOPS GPU outperforms a 1.36 TFLOPS GPU confirms that this benchmark exercises real-world bottlenecks — memory access patterns, compute scheduler overhead, wave occupancy, and driver-side code generation — rather than measuring a synthetic upper bound.

13d. Why the iGPU Wins Despite Lower TFLOPS

Architecture efficiency matters more than shader count. TeraScale 2 uses a VLIW5 (Very Long Instruction Word) design where each "stream processor" is actually five tightly coupled ALUs that must execute in lockstep. If the compiler cannot fill all five slots (a common occurrence for compute shaders with irregular control flow), the vacant slots are wasted. Real-world VLIW5 utilisation in compute workloads is estimated at 50–70%, reducing the HD 5770's effective throughput to roughly 0.7–0.95 TFLOPS.

RDNA 2, by contrast, uses a scalar + SIMD32 design where each compute unit contains two independent SIMD32 units. Every lane executes useful work on every clock — there is no VLIW packing problem. At 2,200 MHz, the 128 shaders deliver nearly their full 0.56 TFLOPS.

The real gap is far smaller than the spec sheet suggests. After accounting for VLIW5 utilisation losses, the effective compute advantage shrinks from the theoretical 2.4× (1.36 vs 0.56 TFLOPS) down to roughly 1.3–1.7× (0.7–0.95 vs 0.56 TFLOPS). The remaining factors below — memory bandwidth parity, driver quality, and compute scheduler maturity — are more than sufficient to close this residual gap and tip the balance in the iGPU's favour.

Compute shader support maturity. The HD 5770 was designed primarily for DirectX 11-era pixel and vertex shading. Its compute shader support (DirectCompute 5.0) was a first-generation implementation with limited occupancy, no asynchronous compute queues, and restricted shared memory bandwidth. RDNA 2 treats compute as a first-class workload with dedicated hardware schedulers, LDS (Local Data Share) bandwidth matched to ALU throughput, and fine-grained wave management.

Driver optimisation. AMD's current Radeon Software Adrenalin Edition drivers for RDNA 2 are actively maintained and optimised. The HD 5770's legacy Crimson Edition drivers (version 16.2.1, Mar 2016) have not received performance updates in over a decade (actually, it‘s real 10 years, now it's Mar 2026). Compute shader code generation for TeraScale 2 was never a priority — these drivers were written when GPU compute was still in its infancy.

Memory bandwidth parity. The HD 5770's theoretical advantage in dedicated GDDR5 is largely neutralised here. Its 76.8 GB/s bandwidth is slightly below the iGPU's ~83 GB/s from dual-channel DDR5-6000 C28. For a bandwidth-sensitive particle simulation, this effectively levels the playing field — or tilts it slightly in the iGPU's favour.

13e. Would the HD 5770 Win in a Gaming Benchmark?

Very likely yes, for traditional 3D rendering workloads. The HD 5770 has 6.25× more shader units, 5× more texture mapping units, and 4× more render output units than the 2-CU iGPU. In a conventional rasterisation pipeline — vertex processing, texture sampling, pixel shading, and blending — these fixed-function resources matter far more than per-CU compute efficiency.

Online gaming benchmarks broadly confirm this: the HD 5770 can run older titles (pre-2015) at low-medium settings, whereas the Ryzen 7600 iGPU struggles to maintain playable frame rates in the same scenarios.

The key insight: This benchmark is a compute-first workload — a particle simulation driven by a compute shader, with a simple instanced rendering pass for visualisation. It exercises the GPU's general-purpose compute pipeline, not its fixed-function rasterisation hardware. The result is a measure of compute shader throughput and scheduling efficiency, where architectural modernity dominates raw shader count.

This makes the benchmark a useful complement to traditional GPU tests. A gaming benchmark tells you how fast a GPU can rasterise triangles; this benchmark tells you how efficiently it can execute general-purpose parallel computation — a workload increasingly relevant to physics simulation, machine learning inference, post-processing, and scientific computing.


Part 5 — Platform and API Issues

14. OpenGL GPU Selection — Platform Limitations

Unlike Vulkan, DirectX 11, and DirectX 12, OpenGL has no standard API for enumerating or selecting a specific GPU on a multi-GPU system. Each of the other backends provides an adapter/device enumeration mechanism:

API GPU Enumeration Per-GPU Selection
Vulkan vkEnumeratePhysicalDevices Create device on any enumerated physical device
DirectX 12 IDXGIFactory::EnumAdapters Pass chosen adapter to D3D12CreateDevice
DirectX 11 IDXGIFactory::EnumAdapters Pass chosen adapter to D3D11CreateDevice
OpenGL None OS/driver decides

Windows

On Windows, the OpenGL context is created by the OS display driver model (WDDM), which assigns the GPU based on system-level configuration. The application has no standard API to override this at runtime.

Available workarounds (limited):

Method Scope Limitation
NvOptimusEnablement export symbol Forces discrete NVIDIA GPU on Optimus laptops Only works on NVIDIA + Intel hybrid laptops; no effect on desktop multi-GPU
AmdPowerXpressRequestHighPerformance export symbol Forces discrete AMD GPU on switchable graphics laptops Same — laptop-only, binary choice (discrete vs integrated)
WGL_NV_gpu_affinity extension Per-GPU context creation Quadro professional cards only — not available on GeForce/consumer GPUs
Windows Graphics Settings panel Per-executable GPU assignment Requires manual user configuration outside the application

On the test system (RTX 5090 + AMD Radeon iGPU desktop), none of the programmatic methods are effective — the only way to force OpenGL onto the integrated GPU is through the Windows Graphics Settings panel.

Linux

Linux provides significantly better OpenGL GPU selection:

Method Scope How
DRI_PRIME=N environment variable Per-process GPU selection (Mesa drivers) DRI_PRIME=1 ./gpu_benchmark
__NV_PRIME_RENDER_OFFLOAD=1 Per-process offload to NVIDIA GPU __NV_PRIME_RENDER_OFFLOAD=1 __GLX_VENDOR_LIBRARY_NAME=nvidia ./gpu_benchmark
EGL_EXT_platform_device Programmatic per-GPU EGLDisplay creation Requires EGL instead of GLX; Mesa 23.3+
EGL_EXT_explicit_device Same, with native windowing support Mesa 23.3+

The application detects Linux at runtime and uses DRI_PRIME to route OpenGL to the user's requested GPU index.

Impact on Benchmarking

This limitation means OpenGL cross-GPU comparisons on Windows require manual configuration, whereas all other backends support interactive GPU selection within the application. On Linux, DRI_PRIME provides equivalent functionality to other backends' built-in GPU selection.


15. DX11 Timestamp Query Failures — Three Distinct Causes

DX11 is the only API in the benchmark where GPU timestamp queries can silently fail to produce results. Vulkan and DX12 always return timestamp values regardless of GPU clock state. DX11, by contrast, uses a D3D11_QUERY_TIMESTAMP_DISJOINT wrapper that can actively refuse to return data.

Three distinct failure modes were observed during testing:

15a. Driver Never Resolves Queries (GetData → S_FALSE indefinitely)

Affected: Windows on ARM virtual machines (SVGA virtual GPU driver).

ID3D11Device::CreateQuery succeeds for both D3D11_QUERY_TIMESTAMP and D3D11_QUERY_TIMESTAMP_DISJOINT, and the application reports timestamps as "enabled". However, ID3D11DeviceContext::GetData for the disjoint query perpetually returns S_FALSE — the result is never ready.

This is a driver limitation: the virtual GPU driver accepts query creation but does not implement the hardware counters needed to resolve them. No application-level workaround exists.

See docs/woa-dx11-timestamp-issue.md for a detailed write-up.

15b. GPU Clock Frequency Instability (Disjoint = TRUE)

Affected: Integrated GPUs under fluctuating load, discrete GPUs during power-state transitions.

The D3D11_QUERY_DATA_TIMESTAMP_DISJOINT structure contains a Disjoint boolean. When TRUE, it signals that the GPU's clock frequency changed during the frame (P-state transition, thermal throttling, power-saving downclock), making the timestamp-to-millisecond conversion unreliable.

The D3D11 specification recommends discarding the entire frame's timing data when Disjoint = TRUE. If the GPU is frequently switching power states — common on integrated GPUs under variable load, or during the first few seconds of a benchmark run while the GPU ramps up — this can result in many consecutive frames with no timing data.

This is not a driver bug. It is a deliberate DX11 design choice to prioritise timestamp accuracy over availability.

Key difference from Vulkan/DX12: Neither Vulkan nor DX12 has a Disjoint concept. Their timestamp queries always return values based on a fixed timestampPeriod / Frequency, even if the GPU clock changes mid-frame. The precision may degrade slightly, but data is never withheld entirely. This is why Vulkan and DX12 report timestamps reliably in scenarios where DX11 reports none.

Mitigation implemented: The application now caches the last known stable frequency (lastGoodFrequency). When Disjoint = TRUE, timestamps are still read and converted using the cached frequency rather than being discarded. This mirrors the behaviour of Vulkan/DX12 — accepting marginally less precise data in exchange for continuous availability.

15c. Query Pipeline Too Shallow (Ring Buffer Depth)

Affected: Slow GPUs (integrated, software renderer) under high particle counts.

If the GPU takes significantly longer than one frame to process submitted work, the application may attempt to read a query result before the GPU has finished writing it. GetData returns S_FALSE because the query genuinely hasn't resolved yet — not because the driver doesn't support it.

Fix: The ring buffer was increased from 4 to 8 slots, and GetData retries were increased to 128 with periodic Sleep(1) yields, giving slow GPUs more time to resolve queries.

Summary Table

Scenario Root Cause Driver Bug? Fix
WoA virtual GPU — never returns data Driver doesn't implement timestamp counters Yes None (graceful fallback to CPU-only timing)
iGPU / dGPU ramp-up — intermittent gaps Disjoint = TRUE during clock transitions No (spec behaviour) Use cached frequency instead of discarding
Slow GPU — first N frames missing Query not resolved before read No (pipeline depth) Deeper ring buffer (8 slots) + retry with Sleep
WARP DX11 — works after warm-up Combination of 6b and 6c No Same mitigations as above

Cross-API Timestamp Mechanism Comparison

Each backend uses its own API's timestamp mechanism — there is no cross-API data sharing. The same GPU executing the same workload produces nearly identical execution times (0.065–0.104 ms on RTX 5090), but the measurement infrastructure differs significantly:

Vulkan DX12 DX11 OpenGL
Write vkCmdWriteTimestamp EndQuery → ID3D12QueryHeap context->End(query) glQueryCounter(GL_TIMESTAMP)
Read vkGetQueryPoolResults with WAIT_BIT ResolveQueryData → readback buffer GetData (CPU polling) glGetQueryObjectui64v
Synchronisation GPU-side wait (guaranteed ready) GPU-side resolve (ordered in command list) CPU polls until S_OK (may never arrive) CPU polls (typically resolves quickly)
Disjoint / clock check None None Required (D3D11_QUERY_TIMESTAMP_DISJOINT) None
Counter frequency Fixed (timestampPeriod), independent of core clock Fixed (GetTimestampFrequency), independent of core clock May vary with GPU core clock Fixed, monotonic counter
Clock-change handling Returns data; timer may reset across submissions† Returns data; stable-clock design, no resets Refuses data if Disjoint = TRUE Returns data, counter is monotonic
First-frame data Yes Yes No (ring buffer warm-up required) Yes (after 1–2 frame delay)

Vulkan caveat: The Vulkan spec notes that power management events (e.g. GPU idle → active transitions) can reset the timestamp counter on some implementations. This affects cross-submission comparisons only — timestamps within the same command buffer are always reliably comparable. The VK_EXT_calibrated_timestamps extension provides monotonic timestamps immune to power events, but is not required for within-frame profiling. In this benchmark, all four timestamps (compute begin/end, render begin/end) are recorded within a single command buffer, so power-state resets do not affect the results.

DX12's stable-clock design: DX12 goes further than Vulkan by explicitly stabilising the GPU clock for timestamp purposes. Two timestamps within the same command list are always comparable, and two timestamps from different command lists are also reliable as long as the GPU did not idle between them. There is no Disjoint equivalent — the API guarantees clock stability by design.

DX11 is the only API that can actively withhold timestamp data based on GPU clock stability. DX12 and Vulkan both use a fixed counter frequency independent of the GPU core clock, so frequency scaling and P-state transitions do not invalidate their results. This makes DX11 the most fragile timestamp implementation from an application developer's perspective, despite the underlying GPU hardware being identical across all backends.


16. OpenGL Compute Shader Performance on AMD GPUs

Observation

OpenGL compute shader performance on AMD GPUs is significantly lower than Vulkan / DX12 / DX11, with older architectures affected most severely.

GPU Architecture OpenGL Compute ms Vulkan Compute ms Ratio
RTX 5090 (reference) Blackwell 0.019 0.019 1.0×
Radeon Graphics (iGPU) RDNA 2 1.489 0.758 2.0×
RX 9070 XT RDNA 4 2.612 0.033 79.2×
RX 6900 XT RDNA 2 2.742 0.184 14.9×
RX 6600 XT RDNA 2 2.719 0.240 11.3×
Vega Frontier Edition Vega (GCN 5) 3.321 0.368 9.0×
RX 580 Polaris (GCN 4) 18.913 0.362 52.3×

On NVIDIA, OpenGL and Vulkan compute times are nearly identical. On AMD, OpenGL compute is 9–52× slower depending on architecture generation.

FPS Impact

GPU OpenGL FPS Vulkan FPS DX11 FPS
RX 9070 XT 253 1,751 1,774
RX 6900 XT 229 2866 4107
RX 6600 XT 180 — —
RX 580 42 783 755

The RX 580's OpenGL score (42 FPS) is lower than the Ryzen 5 7600 CPU-based WARP software renderer running DX11 (44–53 FPS).

Root Cause Analysis

This is a well-documented AMD Windows OpenGL driver limitation, not a code issue:

  • Same code, different results: The identical OpenGL compute path achieves 2062 FPS on RTX 5090 (GPU time 0.019 ms), confirming the shader and API usage are correct.

  • Observed per-dispatch overhead: On all GCN/RDNA GPUs tested in this benchmark, OpenGL glDispatchCompute exhibits a consistent overhead of ~2.7 ms (RDNA 2) to ~18.9 ms (GCN 4), independent of GPU compute capability — the RX 6900 XT and RX 6600 XT show nearly identical compute times despite having 80 vs 32 CUs. While AMD's general OpenGL performance issues are well-documented (see below), specific quantification of per-dispatch compute overhead does not appear to have been published elsewhere — this benchmark may be the first to isolate and measure it. This is demonstrated by the following comparison:

    Metric HD 5770 OpenGL RX 6600 XT OpenGL RX 6600 XT Vulkan
    Compute 1.794 ms 2.719 ms 0.270 ms
    Render 3.018 ms 2.322 ms 0.379 ms
    Total GPU 4.818 ms 5.148 ms 0.649 ms
    FPS 188 180 1239

    The RX 6600 XT's Vulkan compute time (0.270 ms) proves the hardware is 10× faster than what OpenGL reports (2.719 ms). The ~2.7 ms figure appears to be a driver-level overhead floor observed consistently across all GCN/RDNA hardware tested, rather than a reflection of GPU capability. The HD 5770 — a vastly weaker GPU from 2009 — achieves lower OpenGL compute times (1.794 ms) and higher FPS (188 vs 180) than the RX 6600 XT simply because it uses a completely different TeraScale driver stack that does not exhibit this overhead.

  • Known industry issue: Multiple major projects and community reports have documented AMD's OpenGL performance gap on Windows:

  • GCN/Polaris most affected: Since Adrenalin 23.9, AMD moved GCN (Polaris/Vega) to a maintenance-only driver branch with no new performance optimisations. The OpenGL-to-Vulkan translation layer improvements may not have been fully applied to these legacy architectures, explaining the extreme 52× gap on the RX 580.

  • TeraScale counterexample: The HD 5770 (TeraScale 2 / Evergreen) uses a completely different, legacy driver stack and does not exhibit this OpenGL compute overhead. As shown in the table above, it outperforms GCN/RDNA 2 cards in OpenGL despite being vastly inferior hardware. According to Chips and Cheese's architectural analysis, GCN completely rewrote AMD's GPU architecture and driver stack from TeraScale's VLIW design to a scalar SIMD model — meaning the OpenGL driver codebases are entirely separate. The overhead problem was introduced in the GCN/RDNA driver branch, not inherited from TeraScale. Notably, the HD 5770 shows the opposite pattern on DX11: compute + render = 3.87 ms, but total GPU time = 9.15 ms — a 5.3 ms synchronisation overhead between the compute and render stages, likely due to immature compute shader support on this early DX11-era architecture.

Can This Be Optimised at the Application Level?

The OpenGL compute path in this benchmark is already minimal — one glDispatchCompute call and one glMemoryBarrier per frame. There is no room to reduce dispatch frequency, batch operations, or eliminate synchronisation. Alternative approaches such as glDispatchComputeIndirect or persistent mapped buffers do not address the bottleneck, as the overhead originates within the driver's internal dispatch path, not in data transfer or API call volume.

This suggests that when the performance bottleneck is a driver-level fixed cost, application-level optimisation cannot break through the ceiling. The only effective solution is to use an API that avoids this overhead entirely — Vulkan achieves 0.270 ms for the same compute workload that takes 2.719 ms through OpenGL on the same hardware (RX 6600 XT), a 10× improvement with identical shader logic.

Why Did Minecraft See 79–92% Improvement but This Benchmark Did Not?

AMD's OpenGL-to-Vulkan translation layer excels at optimising high-volume rendering workloads. Minecraft Java Edition issues thousands of draw calls per frame (blocks, entities, particles, UI), each carrying CPU-side overhead for state changes and submission. The translation layer batches these into efficient Vulkan command buffers, dramatically reducing the per-call cost — e.g., 1000 draw calls × 0.1 ms overhead each = 100 ms, batched down to a few grouped submissions at ~10 ms total.

This benchmark has the opposite profile: one single glDispatchCompute call per frame with a ~2.7 ms observed driver overhead. The translation layer's batching strategy cannot help here — there is nothing to batch. The overhead appears to be a per-dispatch fixed cost within the driver, not an accumulation of many small costs that can be amortised. This is why Minecraft saw up to 92% improvement while our compute workload on the same latest drivers (Adrenalin 26.3.1) shows no meaningful change.

Note: The Khronos Community Forums contain a thread on glDispatchCompute calling overhead, but that discussion reports ~0.2 ms overhead on older NVIDIA GPUs (GTX 560/470), not AMD. The ~2.7 ms overhead observed in this benchmark on AMD GCN/RDNA hardware is an order of magnitude larger and does not appear to have been specifically documented elsewhere.

GTX 970 (Maxwell): DX11 Compute–Render Synchronisation Overhead

The GTX 970 exhibits a striking reversal in API performance ranking compared to the RTX 5090:

API Compute Render Compute+Render Total GPU FPS Bottleneck
Vulkan 0.434 ms 0.663 ms 1.097 ms 1.098 ms 718.8 Balanced
OpenGL 0.431 ms 0.791 ms 1.222 ms 1.226 ms 642.2 Balanced
DX12 0.443 ms 0.535 ms 0.978 ms 0.977 ms 291.1 CPU-bound
DX11 0.435 ms 0.873 ms 1.308 ms 3.355 ms 280.3 GPU-bound

On the RTX 5090, DX11 is the fastest API (7736 FPS). On the GTX 970, it is the slowest (280 FPS). The cause is visible in the numbers:

  • DX11: Compute + render sum to only 1.308 ms, but total GPU time is 3.355 ms — a ~2 ms synchronisation overhead between the compute and render stages. Maxwell's DX11 driver appears to insert a costly pipeline flush/barrier when transitioning from compute dispatch to draw calls. The RTX 5090 (Blackwell) shows no such overhead (compute + render ≈ total GPU time).
  • DX12: GPU time is actually the fastest (0.977 ms), but FPS is only 291.1 — a massive 2.5 ms CPU overhead (frame time 3.435 ms − GPU 0.977 ms). This reflects the high CPU-side cost of DX12 command recording on an older driver/architecture combination.
  • Vulkan: Best overall balance — GPU time 1.098 ms with only 0.293 ms CPU overhead. The explicit API model with pre-recorded command buffers works well even on older hardware.
  • OpenGL: NVIDIA's OpenGL driver performs well (unlike AMD), with GPU time 1.226 ms and minimal CPU overhead.

This demonstrates that API performance rankings are not universal — they depend on GPU architecture and driver maturity. DX11's implicit driver model excels on modern NVIDIA hardware (where the driver has been refined over a decade) but introduces overhead on older architectures where compute–render transitions were not as optimised. The RX 9070 XT (RDNA 4) shows yet another pattern: DX11 ≈ Vulkan ≈ DX12, with only OpenGL significantly behind — AMD's DX11 advantage over explicit APIs is negligible compared to NVIDIA's.

Direct3D 10-era GPUs: GT 120 / 9500 GT Downlevel Path

The earlier version of this report incorrectly concluded that Feature Level 10_0 could not execute compute shaders at all. The observed failure was real, but its cause was in this benchmark: the DX11 backend created an FL10_0 device and then unconditionally compiled cs_5_0, vs_5_0, and ps_5_0. Naturally, CreateComputeShader rejected that Shader Model 5 bytecode on an SM4 device.

Direct3D 11 exposes an optional, genuine DirectCompute 4.x path for Direct3D 10.0/10.1 hardware. Microsoft requires applications to query D3D11_FEATURE_D3D10_X_HARDWARE_OPTIONS and specifically ComputeShaders_Plus_RawAndStructuredBuffers_Via_Shader_4_x; support cannot be assumed from the model name alone. See Microsoft's downlevel compute documentation and the feature-query structure.

The current backend now:

  • stores the device's actual feature level;
  • queries the optional DirectCompute 4.x capability on FL10 hardware;
  • compiles SM4.0 compute/vertex/pixel profiles on FL10 and SM5.0 on FL11;
  • completely skips compute shaders and UAV particle buffers for fragment-only workloads, so those tests remain available even when the optional compute bit is absent;
  • rejects FP64 SynthPeak and oversized dispatches cleanly on the downlevel path;
  • records feature level, shader model, and DirectCompute availability with the saved driver metadata.

SM4 downlevel compute is constrained but real: only one compute UAV can be bound, typed UAVs are unavailable, thread-group shared memory is limited to 16 KiB, and a group is limited to 768 threads. The original Stream shader uses 256 threads and one RWStructuredBuffer, so it fits this contract. N-body now uses SV_GroupIndex, allowing its small configurations to compile for SM4; SynthPeak FP32/INT32 also compiles, while FP16/FP64 remain excluded.

Era DirectX / FL Shader Model Relevant capability
Advanced raster DX9.0c SM 3.0 VS/PS, dynamic branching; no compute/UAV
Unified raster DX10 / FL10_0 SM 4.0 VS/PS/GS and optional DirectCompute 4.x through the D3D11 runtime
Refined unified DX10.1 / FL10_1 SM 4.1 SM4.1 raster additions and optional DirectCompute 4.x
Full DX11 compute DX11 / FL11_0 SM 5.0 Standard compute, richer UAV/TGSM/atomic/tessellation feature set
Explicit APIs DX12 / Vulkan SM 5.1+ / SPIR-V Explicit command recording, queues, and modern resource control

This does not mean every workload is safe to launch at its modern default on a GT 120. First acceptance should use Stream/Particle Light or Medium and small N-body/Render3D configurations. Large SynthPeak loops, GPU Burn/Stress, and Volumetric need a separately versioned legacy_sm4 calibration to avoid Windows TDRs. FP16, FP64, Vulkan, DX12, OpenGL 4.3, Fluid, and Cinematic Liquid are not part of the GT 120 contract.

Why a new DX9 backend is not the GT 120 solution

DX9 could be implemented as a separate raster backend, but it has no compute shader/UAV path and would require a new D3D9 device, presentation, timestamp, resource, and SM3 shader implementation. Its results could not share the Stream, N-body, or SynthPeak contracts. RenderDoc also does not support D3D9 capture, so it would break this project's fixed fifth-second capture workflow. Because the GT 120 already exposes FL10_0, D3D11 downlevel is both the more capable and the more comparable route.

The code and all SM4 profiles have been compiled on the development machine, but the optional DirectCompute bit, driver timing behavior, safe calibration, and fifth-second RenderDoc capture still require validation on the physical GT 120 before any result is called formal.

Shader Model, Shading Languages, and the Rendering Pipeline

Three concepts are often conflated but serve distinct roles:

  • Rendering pipeline — the GPU's processing stages (vertex → rasterisation → fragment → output, plus optional stages like geometry, tessellation, and compute). This defines what stages exist and how data flows between them.
  • Shader Model (SM) — a hardware capability specification defining what code can run at each programmable stage. Higher SM versions unlock more instructions, longer programs, new memory access patterns, and new pipeline stages.
  • Shading languages — the programming languages developers use to write shader code for each stage.

These three evolve together: a new pipeline stage (e.g., compute shader) requires new hardware capability (SM 5.0), which is then exposed through shading language features (HLSL [numthreads], GLSL layout(local_size_x)).

Shading Languages and Compilation

Each graphics API defines its own shading language:

Language API Compilation Target
HLSL (High-Level Shading Language) DirectX 9–12 fxc (SM 2.0–5.0) / dxc (SM 6.0+) DXBC / DXIL bytecode
GLSL (OpenGL Shading Language) OpenGL / Vulkan glslang / glslc OpenGL: driver compiles at runtime; Vulkan: pre-compiled to SPIR-V
MSL (Metal Shading Language) Metal Metal compiler Apple IR
SPIR-V Vulkan Intermediate representation Consumed by Vulkan driver, compiled to GPU-native ISA

The compilation pipeline in this benchmark:

Vulkan:   GLSL (.comp/.vert/.frag)  ──→  glslc  ──→  SPIR-V (.spv)  ──→  Vulkan driver  ──→  GPU ISA
DX12:     HLSL (.hlsl)              ──→  fxc    ──→  DXBC bytecode   ──→  DX12 driver    ──→  GPU ISA
DX11:     HLSL (.hlsl)              ──→  fxc    ──→  DXBC bytecode   ──→  DX11 driver    ──→  GPU ISA
OpenGL:   GLSL (embedded strings)   ──→  driver compiles at runtime  ──→  GPU ISA

Despite using different languages, all backends compile down to the same GPU instruction set (ISA) for a given GPU. The shader logic is equivalent across all four backends — position update in the compute shader, point-sprite rendering in the vertex/fragment shaders. The only differences are syntax and API-specific boilerplate.

This is why cross-API comparisons in this benchmark are meaningful: the same algorithm runs through different API/driver paths to the same hardware, isolating the API and driver overhead from the shader workload itself.

Shader Model and Feature Level Mapping

The Shader Model version determines what a GPU can do, but different APIs expose this through different naming:

SM DX Feature Level OpenGL Version GLSL Version Key Addition
SM 2.0 9_1 – 9_3 2.1 120 Basic VS/PS, FP32
SM 3.0 — 3.0 130 Dynamic branching
SM 4.0 10_0 3.3 330 Unified shaders, geometry shader, integer ops; optional DirectCompute 4.x through D3D11
SM 4.1 10_1 — — Gather4/MSAA additions; optional DirectCompute 4.x through D3D11
SM 5.0 11_0 4.3 430 Standard full compute/UAV feature set and tessellation
SM 5.1 11_1 / 12_0 4.5+ 450 Bindless-style resource indexing
SM 6.0+ 12_0+ — (Vulkan SPIR-V) — Wave intrinsics, ray tracing, mesh shaders

The modern cross-API comparison matrix still starts at SM 5.0 / Feature Level 11_0 / OpenGL 4.3. A new, explicitly separated DX11-downlevel contract can include FL10_0/SM4 hardware when its driver exposes DirectCompute 4.x. Those results must record the feature level and shader profile and must not be mixed with Vulkan/DX12/OpenGL 4.3 coverage claims.

Conclusion

These results quantitatively demonstrate that modern APIs (Vulkan, DX12) are essential for realising the full compute potential of AMD hardware. The OpenGL compute path carries significant driver overhead on AMD, particularly on legacy architectures, reinforcing the industry trend towards explicit, low-overhead graphics APIs. Choosing the right API is more impactful than optimising application code when driver-level overhead dominates.

The RX 9070 XT (RDNA 4) is the most extreme example: OpenGL compute takes 2.612 ms vs Vulkan's 0.033 ms — a 79× penalty — the largest ratio observed in any GPU tested. Despite being AMD's newest architecture, the OpenGL compute dispatch overhead has not improved from RDNA 2 levels (~2.7 ms), confirming that AMD's driver team has deprioritised OpenGL compute optimisation.

Additionally, API performance rankings are architecture-dependent: DX11 leads on modern NVIDIA GPUs but falls behind Vulkan on older Maxwell hardware due to compute–render synchronisation costs. On the RX 9070 XT, DX11 and Vulkan are nearly tied (1,774 vs 1,751 FPS), reflecting AMD's less mature DX11 optimisation path. This underscores that no single API is universally optimal — the best choice depends on the target hardware generation and driver maturity.


17. Dual Identical GPU Behaviour (Mac Pro 2013 — 2× FirePro D700)

The Mac Pro (Late 2013) contains two identical AMD FirePro D700 GPUs (GCN 1.0, Tahiti XT). Two concepts must be kept separate on this machine: selecting one adapter independently, and making both adapters collaborate on one benchmark. Earlier report revisions conflated them.

AFR, SFR, CrossFire and SLI — brief definitions

Before the D700 measurements, four terms are easy to mix up. AFR and SFR are how two GPUs share work inside one application. CrossFire (AMD) and SLI (NVIDIA) are vendor multi-GPU brands / driver frameworks that historically packaged those modes for games—usually as an opaque driver profile rather than an app-controlled API contract.

Term What it means How work is split Typical cost / risk
AFR (Alternate Frame Rendering) GPUs take turns rendering whole frames (GPU0 → frame N, GPU1 → frame N+1, …) Frame-level parallelism; both GPUs can overlap if the queue is deep enough Extra frames-in-flight; higher input latency; bad for strongly stateful frame-to-frame dependencies unless state is copied
SFR (Split Frame Rendering) Both GPUs render parts of the same frame (commonly left/right halves, or other screen partitions), then one present path composites Intra-frame data parallelism; one logical image Cross-GPU copy/composite every frame; can lose to AFR when the transfer cost dominates light shaders
CrossFire AMD’s multi-GPU product/driver branding (later “mGPU” naming varied by generation) Historically often AFR via driver profiles; some eras also SFR / hybrid modes App may get no control and no honest utilization signal; profile-dependent and largely retired for modern DX12/Vulkan games
SLI NVIDIA’s multi-GPU product/driver branding Same idea family as CrossFire—mostly driver-managed AFR for DX11-era titles Same opacity problem; modern NVIDIA dual-GPU gaming support is effectively discontinued

Key distinctions for this benchmark:

  1. Mode vs brand. AFR/SFR describe the scheduling. CrossFire/SLI name the vendor stack that might implement a mode for you. Saying “CrossFire is on” does not prove the app is doing explicit AFR or SFR—only that the driver claims multi-GPU participation.
  2. Explicit engine path vs implicit driver path. Mangekyo’s validated collaboration is application-controlled DX12 linked-adapter work (NodeMask, per-node queues/resources, --multi-gpu afr|sfr). That is not the same as hoping DX11 CrossFire/SLI profiles accelerate a custom engine. On this D700 FireGL stack, AMD AGS still saw two physical cards but returned crossfireAPI=0 / one active GPU for both explicit and driver AFR requests—so classic CrossFire was unavailable even when two adapters existed.
  3. Why Plasma can use AFR, while Particle prefers SFR. Plasma/GPU Burn has no persistent simulation state between frames, so alternating whole frames is meaningful and measured ~2× on DX12 AFR (with ≥4 frames in flight, RenderDoc off). Original Particle’s frame N+1 depends on frame N particle state; AFR would require a full cross-node state copy every frame or would change the test semantics. The product direction for Particle is therefore fixed-total-count SFR/data-parallel split of one frame, not AFR.
  4. SLI is not “validated” here. The DX12 linked-node code is vendor-neutral in the sense Microsoft’s linked-GPU samples describe homogeneous CrossFire/SLI topologies, but this report only accepts measured D700 DX12 AFR/SFR results. No NVIDIA SLI system has been validated; the GUI must not claim SLI support.

In short: AFR = alternate whole frames; SFR = split one frame; CrossFire/SLI = vendor multi-GPU brands that may hide either mode behind the driver. This project only scores modes it controls and measures.

Addressability and collaboration

API Independent adapter selection Explicit two-GPU collaboration D700 result
Vulkan Two VkPhysicalDevice objects; GPU #2 can be selected directly Physical-device-group masks Functional, but the only graphics-capable family exposes one VkQueue, so tested AFR submissions serialize and do not scale
DirectX 12 Driver-dependent: 25.20.14020.10001 exposed two distinct DXGI LUIDs but routed both to the primary GPU; 27.20.14540.15002 exposes one linked adapter Linked-adapter device with two explicit nodes and NodeMask control Working AFR (~2×) and workload-dependent SFR (1.47× at 128 steps)
DirectX 11 The old driver accepted two DXGI LUIDs but routed both to the primary GPU; the current topology exposes one logical adapter and AGS still sees both physical D700s Standard D3D11 has no node control; AGS can request explicit or driver-managed CrossFire when the driver exposes it Windowed implicit AFR did not scale; AGS returned crossfireAPI=0 and one active GPU for both AFR modes
OpenGL The AMD WGL extension enumerates both D700 GPU IDs WGL_AMD_gpu_association would permit per-GPU off-screen contexts and cross-context framebuffer blit Driver exposes the interface but refuses to create an associated context for GPU #2; explicit SFR is unavailable

DXGI LUID matching remains necessary because factory instances can reorder adapters, and VendorId + DeviceId + SubSysId is not a unique identity. The archived March run on 25.20.14020.10001 recorded two distinct D700 LUIDs and successfully created DX11/DX12 devices through both handles. Task Manager nevertheless showed both D3D runs executing on the primary physical GPU, so those are valid API runs but not independent second-card measurements. After reinstalling 27.20.14540.15002, the current probe sees one D700 DXGI LUID; ID3D12Device::GetNodeCount() exposes its two physical nodes. D3D11 cannot address those nodes independently, but D3D12 can: the benchmark now expands a linked adapter into selectable node rows and binds all ordinary single-GPU objects and resources to the selected node mask.

A July 21 run-all check on the current driver exposed a probe-merging and routing bug: Vulkan correctly returned two VkPhysicalDevice objects, but the synthetic second row copied the first row's DXGI identity while the ordinary DX12 backend left NodeMask=0. Consequently both labelled DX12 runs executed on linked-adapter node 0; the two 65K Particle results (82.610 and 82.723 GB/s) were not independent card measurements. The corrected probe derives two rows from GetNodeCount() and correlates the two Vulkan devices with those rows. The verified capability table is now D700 #1 = Vulkan/DX12/DX11/OpenGL and D700 #2 = Vulkan/DX12. For ordinary DX12, the command queue, ResizeBuffers1 placement, command lists, root signatures, PSOs, descriptor/query heaps, timestamps, and particle/fluid resources all use the row's explicit mask (0x1 or 0x2). A temporary 4M Particle validation produced 122.8 FPS / 78.66 GB/s on node 0 and 120.1 FPS / 80.14 GB/s on node 1. Each result persists dx12Node, dx12NodeMask, and dx12LinkedNodes in workloadConfig. This correction does not retroactively assert that the old driver's two LUIDs were fabricated; it records that the old physical routing and the new linked-node topology are different cases.

The Windows GUI now exposes Multi-GPU: Off / AFR / SFR with an adjacent information icon. It is enabled only for a Custom run whose sole API is DX12 and whose workload is Plasma; selecting AFR or SFR forwards the matching --multi-gpu argument and disables RenderDoc, while changing to an unsupported combination resets the mode to Off. The code is not AMD-specific: a NVIDIA SLI configuration could use the same path only when its driver presents the pair as one DX12 linked adapter with GetNodeCount() >= 2 and the required cross-node sharing support. Microsoft's official D3D12 linked-GPU sample explicitly describes the linked homogeneous case as CrossFire/SLI, but this benchmark has not yet validated an NVIDIA system and therefore does not claim SLI support.

History now renders the saved execution contract explicitly in a Mode column: Single, Headless, AFR ×2, or SFR ×2. Legacy DX11/Vulkan AFR experiments are labelled unverified/experimental from their persisted control tags. The headless boolean was already serialized, but the generic comparison grouping previously omitted it and could place windowed and headless stream_v1 rows together. Comparison groups now include execution mode, and pairwise score deltas require matching headless state.

The final DX11 probe used the official AMD AGS 6.3.1 library rather than inferring CrossFire activity from Task Manager. agsInitialize enumerated both FirePro D700s. agsDriverExtensionsDX11_CreateDevice also succeeded in both AGS_CROSSFIRE_MODE_EXPLICIT_AFR and AGS_CROSSFIRE_MODE_DRIVER_AFR, but both returned extensionsSupported.crossfireAPI=0 and crossfireGPUCount=1. Thus this FireGL driver exposes two physical adapters but does not activate or expose DX11 CrossFire for the application; there is no useful AGS integration to retain in the benchmark. AMD defines explicit AFR as the no-profile path and crossfireGPUCount as the number of GPUs active for the app in the official AGS DX11 API.

The OpenGL SFR probe used the only relevant explicit Windows mechanism, WGL_AMD_gpu_association. FireGL 20.45.40.15 reports two IDs and maps the visible context to ID 1, but ID 2 returns no renderer string. Both wglCreateAssociatedContextAttribsAMD for a 4.3 core context and legacy wglCreateAssociatedContextAMD return NULL, including a retry with the visible context unbound; the ICD also leaves GetLastError at zero. Consequently the second D700 cannot receive OpenGL commands. The planned left/right rendering plus wglBlitContextFramebufferAMD composition cannot begin, so the prototype and CLI enablement were reverted rather than falling back to single-GPU rendering. The intended mechanism and its requirement for a valid GPU-associated context are defined by the Khronos WGL_AMD_gpu_association specification.

Plasma AFR acceptance results (2026-07-21)

The first correct AFR vertical slice is limited to the windowed Plasma/GPU Burn workload (--multi-gpu afr). It alternates complete frames between DX12 node masks 0x1 and 0x2 and uses a separate queue and frame resources for each node.

Backend / configuration 16 shader steps 128 shader steps Interpretation
DX12 single GPU ~113 FPS ~19 FPS Baseline
DX12 AFR, 2 frames in flight ~109 FPS — Too little queue depth; each node has only one reusable slot
DX12 AFR, 4 frames in flight 222–230 FPS 39–40 FPS 1.96–2.05×; both nodes overlap
Vulkan single / device-group AFR ~102 / ~104 FPS ~17 / ~17 FPS No material gain; shared graphics VkQueue serializes work
DX11 single / implicit-driver AFR request ~120 / ~120 FPS ~20 / ~20 FPS No gain; driver-managed physical split unverified

The program therefore forces at least four frames in flight for AFR. Utilisation analysis is reported as GPU-equivalent work across the two devices: the validated DX12 run reached approximately 200% aggregate (about 100% per GPU), whereas Vulkan and DX11 remained near 100% aggregate.

Additional Vulkan isolation tests moved vkQueuePresentKHR from queue family 0 to the independently present-capable transfer family 2. This removed presentation operations from the sole graphics queue without changing the result: 16-step AFR remained about 103 FPS, while the 128-step single/AFR pair remained 17/17 FPS with approximately 100% aggregate GPU-equivalent utilisation. A paired-acquire experiment intended to enqueue both device-mask submissions before either present instead blocked indefinitely in AMD 20.45.40.15's second vkAcquireNextImage2KHR call on the LOCAL-only swapchain and was reverted.

Both D700 multi-instance device-local heaps expose peer access as VK_PEER_MEMORY_FEATURE_COPY_DST_BIT only in either direction—no peer copy-source or generic reads. A secondary GPU can therefore write a destination allocated for the primary GPU, but the primary cannot directly pull from secondary-local output. The remaining single-VkDevice experiment used the compute-only family's two queues and an R8G8B8A8_UNORM storage-capable swapchain: alternate device masks dispatched both Plasma layers directly into each GPU's LOCAL swapchain image. Runtime capability checks, compilation and a RenderDoc-disabled three-second run all succeeded, including clean workload completion. The result was nevertheless only about 119 FPS / 8.18 ms at 16 steps and approximately 98% aggregate GPU-equivalent utilisation (49% average per GPU), proving that the driver still did not overlap alternate frames. Because the compute shader also changes the pipeline contract relative to fragment Plasma, this small throughput difference is not evidence of AFR scaling. The implementation, shader variant and separate score identity were reverted in full. Only a future, separately scoped two-VkDevice external-memory/external-semaphore design remains unexplored. The peer-memory flag meanings are defined by Khronos.

RenderDoc constraint

RenderDoc is not disabled globally. Single-GPU runs still use the normal 15-second run and fifth-second capture. It is automatically disabled for multi-GPU AFR and SFR because merely injecting RenderDoc on this old D700 stack produced false ~2000 FPS output, missing GPU timestamps and a long queue drain at process exit. SFR has not independently proved capture-safe. Multi-GPU scoring must use native timing plus GPUView/PIX-style diagnostics; a one-node reproduction may still be captured separately for shader debugging.

Particle AFR and SFR direction

DX12 can alternate particle frames at the API level, but Original Particle is stateful: frame N+1 consumes the positions and velocities produced by frame N. Copying that full state across nodes every frame would add synchronization and transfer cost, while keeping one independent state per node changes the simulation and is not a valid acceleration. DX11 offers only implicit driver AFR and cannot solve this dependency explicitly. Consequently, particle AFR is not a product path.

The meaningful two-GPU particle design is SFR/data parallelism: keep one fixed total particle count, split the particle range between both GPUs, simulate and draw both partitions for the same frame, then composite once. That measures one accelerated test rather than two tests run side by side. Plasma is the lower-risk SFR prototype because it has no persistent simulation state: render the left and right halves into node-local targets, copy one half to the presenting node, then compose and present. Complexity is moderate rather than trivial—cross-node resource visibility, two fences, copy/composition and failure fallback must all be handled.

The version 0.2.2 runtime probe reports D3D12_CROSS_NODE_SHARING_TIER_1 on the D700 linked adapter. This is sufficient for explicit cross-node CopyBufferRegion, CopyTextureRegion and CopyResource operations when the shared resource is the copy destination, so a Plasma half-frame copy/composite prototype is technically viable. It does not promise free peer bandwidth or cross-node render-target use; the copy cost must be measured before SFR is accepted. See Microsoft's D3D12_CROSS_NODE_SHARING_TIER contract.

DX12 SFR implementation and D700 result

The prototype is now implemented as --multi-gpu sfr / --sfr. Both nodes use the same frame time and shader contract. Node 0 shades the left half directly into the swap-chain buffer; node 1 shades the right half into a node-local full-size render target so SV_Position and UV values remain identical to the single-GPU image. Node 1 then copies only the 640x720 right region into a 1,843,200-byte node-0-owned shared buffer. A cross-node fence lets node 0 consume that buffer, copy it into the right side of the back buffer, and perform the only Present. This is one cooperative frame, not two benchmark instances.

The old D700 driver rejects a cross-adapter committed resource with E_INVALIDARG; using the explicit shared-heap plus placed-resource pattern from Microsoft's heterogeneous multi-adapter sample succeeds. Two and four frames in flight produce the same result, ruling out a shallow frame ring. Short isolated comparisons on driver 27.20.14540.15002 were:

Plasma fixed steps Single DX12 DX12 SFR Scaling Interpretation
16 112 FPS / 8.17 ms 69 FPS / 14.33 ms 0.62x Tier-1 transfer/composition dominates
128 19 FPS / 49.75 ms 28 FPS / 35.11 ms 1.47x half-frame shading repays the fixed copy cost

SFR therefore works and accelerates sufficiently heavy Plasma frames, but it is not a universal optimisation and remains slower than the measured 128-step AFR result of about 39 FPS. Results are isolated under ..._sfr2; RenderDoc remains disabled for SFR until linked-node capture is independently trustworthy. Microsoft's Tier-1 contract permits only cross-node copies with the shared resource as destination, which is why the secondary GPU cannot render directly into the presenting resource. See the official cross-node sharing tiers and multi-adapter copy example.


Part 6 — Conclusion

18. Summary

Observation Explanation
DX11 > DX12 > Vulkan > OpenGL in FPS (hardware GPU, simple workload) Per-frame CPU overhead: DX11 (0.008 ms) < DX12 (0.088 ms) < Vulkan (0.183 ms) < OpenGL (0.322 ms)
All four APIs have similar GPU execution times (0.065–0.104 ms) The GPU-side workload is identical; only CPU-side driver overhead differs
OpenGL has fastest GPU time but lowest FPS WGL swap path and state-machine validation overhead dominate CPU time
DX12 > DX11 in FPS (WARP software renderer) DX11's implicit driver layer is pure overhead when there is no GPU hardware to optimise for
DX12 has lowest GPU time on hardware Slightly more efficient GPU command scheduling, but CPU overhead negates the advantage at low complexity
dGPU ~37× faster than iGPU Bandwidth-bound workload; ratio matches memory bandwidth difference
Device-local 35× faster than host-visible on dGPU PCIe round-trips dominate compute shader memory access patterns
Host-visible = Device-local on iGPU Unified memory architecture — no PCIe hop
WARP ~80× slower than hardware GPU Expected for CPU-based software rasterisation
WARP can appear as a Vulkan device via Dozen Mesa Dozen (Vulkan→D3D12) wraps WARP; only visible when no hardware Vulkan ICD is present
OpenGL cannot select GPU on Windows No standard API; OS assigns GPU. Linux provides DRI_PRIME as a workaround
DX11 timestamps fail where Vulkan/DX12 succeed DX11's Disjoint flag discards data on GPU clock changes; Vulkan/DX12 have no such mechanism
RDNA 2 iGPU (0.56 TFLOPS) beats HD 5770 (1.36 TFLOPS) by 42–71% VLIW5 utilisation losses, immature compute scheduler, and stale drivers reduce TeraScale 2's effective throughput well below its theoretical peak
TFLOPS is a poor predictor of compute shader performance Architectural efficiency (SIMD vs VLIW), driver maturity, and compute scheduler design dominate raw ALU count
HD 5770 would likely win a gaming benchmark Traditional rasterisation relies on fixed-function units (TMUs, ROPs) where the HD 5770 has 4–6× more hardware than the iGPU
Fast GPUs report inflated "render time" in windowed mode Vulkan's COLOR_ATTACHMENT_OUTPUT_BIT timestamp includes swapchain semaphore wait; on fast GPUs (9070 XT), ~90% of reported render time is idle wait
Increasing swapchain BufferCount does not fix timestamp pollution The semaphore wait is inherent to the Vulkan presentation model, independent of buffer pool size
Headless mode achieves 10× FPS over windowed mode Removing swapchain/render/present eliminates presentation overhead; all APIs converge to ~0.034 ms compute time
DX11 headless requires workarounds for timestamp queries Without Present() as a frame boundary, DX11's timestamp pipeline produces garbage values; sanity filtering discards ~3–4% of samples
OpenGL headless requires periodic glFinish() on AMD AMD's driver doesn't actively process commands for hidden windows; periodic full sync every 16 frames restores timestamp availability
3DMark Unlimited ≠ headless compute 3DMark Unlimited renders offscreen (full pipeline); this benchmark's headless mode skips rendering entirely (compute only)
RX 9070 XT (RDNA 4) has 4.1× better per-CU compute than RX 6600 XT (RDNA 2) 2× the CU count (64 vs 32), 1.15× higher clocks; the remaining ~1.8× is architectural improvement in scheduler, cache, and driver codegen
9070 XT outperforms 80-CU RX 6900 XT in headless compute (21K vs 16K FPS) RDNA 4's per-CU efficiency (~1.7× higher) overcomes the 1.25× CU disadvantage; on a per-dollar basis, 9070 XT delivers 1.9× the FPS/$ of RTX 5090
16M particles reveals primitive throughput bottleneck RX 6900 XT (218 FPS) beats 9070 XT (111 FPS) at 16M — 80 CUs provide 1.25× more rasterisation hardware; RTX 5090 (612 FPS) leads due to 170 SMs
9070 XT compute scales 59× for 16× particles (1M → 16M) Super-linear scaling due to GPU underutilisation at 1M; at 16M the GPU is fully saturated
9070 XT headless achieves 21K FPS across all APIs All APIs converge to 0.034 ms compute; windowed mode's 1.7K FPS is 12× slower due to presentation overhead

These results demonstrate that API overhead, memory placement, and hardware architecture all significantly affect GPU compute performance — and that the optimal configuration depends on workload complexity and hardware topology.


Appendices

Appendix: ATI, AMD, and Qualcomm Adreno — A Shared GPU Heritage

The Adreno 640 tested in this benchmark shares a direct lineage with the Radeon GPUs it is compared against. This section traces the corporate and technical connections from ATI Technologies through AMD to Qualcomm's Adreno mobile GPU division.

Corporate Lineage

Year Event
1985 Array Technology Inc. (ATI) founded in Markham, Ontario, Canada
1987 First product: EGA Wonder — ISA graphics card for IBM PCs
1991 ATI enters the dedicated 2D accelerator market (Mach 8, Mach 32)
1996 3D Rage — ATI's first 3D-capable GPU. Competed with 3dfx Voodoo and S3 ViRGE
2000 Radeon DDR (R100) — ATI's first GPU under the Radeon brand. Hardware T&L, competed with NVIDIA GeForce 2
2002 Radeon 9700 Pro (R300) — first DirectX 9 GPU, SM 2.0, outperformed GeForce FX. Widely considered ATI's finest moment
2006 AMD acquires ATI Technologies for $5.4 billion. ATI's GPU division becomes AMD Graphics
2008 AMD sells its mobile GPU division (Imageon) to Qualcomm for $65 million
2009 Qualcomm renames Imageon to Adreno (anagram of "Radeon"). First product: Adreno 200 in Snapdragon QSD8250
2013 AMD rebrands consumer GPUs from "Radeon HD" to "Radeon R" series, then later "Radeon RX"
2017 AMD launches Vega architecture (GCN 5). "ATI" name fully phased out from all products
2020 AMD launches RDNA 2 (RX 6000 series). Qualcomm launches Adreno 660 (Snapdragon 888)
2024 Qualcomm launches Snapdragon X Elite with Adreno X1 GPU for Windows on ARM laptops
2025 AMD launches RDNA 4 (RX 9070 XT). Qualcomm's Adreno GPUs power the majority of Android devices and Windows on ARM PCs

Technical Connection: Imageon → Adreno

ATI's Imageon was a low-power mobile GPU line designed for handheld devices and embedded systems (PDAs, early smartphones). When AMD acquired ATI in 2006, Imageon became part of AMD's portfolio but was considered non-core — AMD's focus was on discrete desktop/laptop GPUs (Radeon) and professional workstation GPUs (FirePro).

In 2008, AMD divested the Imageon mobile GPU division to Qualcomm for $65 million — a fraction of the $5.4 billion AMD paid for all of ATI. Qualcomm integrated Imageon into its Snapdragon SoC platform and renamed it Adreno — an anagram of "Radeon" that preserves the ATI heritage while establishing a distinct brand.

Adreno's architecture has since diverged significantly from Radeon. By the Adreno 600 series (2018), the GPU shares no meaningful silicon design with contemporary Radeon GPUs — the instruction set, memory hierarchy, shader core layout, and driver stack are entirely Qualcomm-designed. However, foundational concepts from ATI's Imageon era (TBR, tile-based rendering optimisations, unified shader architecture for mobile power budgets) persist in Adreno's design philosophy.

TBR vs IMR — and the modern reality. The rendering architecture inherited from Imageon is Tile-Based Rendering (TBR): the screen is divided into small tiles (e.g. 16×16 or 32×32 pixels), and each tile is rendered entirely within fast on-chip tile memory before being written to DRAM — minimising bandwidth-hungry framebuffer read/write traffic. Traditional desktop GPUs use Immediate Mode Rendering (IMR): triangles are rasterised and written to the framebuffer in submission order, relying on high memory bandwidth rather than on-chip buffering. In practice, however, no modern GPU is purely TBR or purely IMR. Mobile GPUs like Adreno and Apple's M-series use TBDR (Tile-Based Deferred Rendering), which adds deferred visibility testing (Hidden Surface Removal) per tile to eliminate overdraw before shading. Meanwhile, desktop GPUs have adopted tile-like techniques: NVIDIA's Maxwell and later use a tiling rasteriser that groups fragments into screen-space tiles for more efficient L2 cache usage; AMD's RDNA series introduced binning passes (a form of tiling) to reduce bandwidth — the Infinity Cache discussed throughout this report is particularly effective when combined with RDNA's binning, as tiles that fit in cache avoid VRAM round-trips entirely. The distinction today is a spectrum: mobile GPUs lean tile-heavy with optional immediate fallback for complex geometry, while desktop GPUs are fundamentally immediate-mode but borrow tiling techniques for bandwidth efficiency.

GPUs Tested in This Benchmark — Family Tree

ATI Technologies (1985)
├── Radeon R100 (2000)
│   └── Radeon 9700 (R300, 2002) — first DX9 GPU
│       └── Radeon X1800 (R520, 2005) — last ATI-only design
│
├── Imageon (mobile GPU line, 2002–2008)
│   └── [Sold to Qualcomm, 2008]
│       └── Adreno 200 (2009) — renamed from Imageon
│           └── Adreno 3xx/4xx/5xx/6xx
│               └── Adreno 640 (2019) ← TESTED: Xiaomi Pad 5 / Snapdragon 860
│                   └── Adreno X1 (2024) — Snapdragon X Elite for WoA
│
└── [AMD acquires ATI, 2006]
    └── AMD Radeon
        ├── HD 5770 (TeraScale 2, 2009) ← TESTED
        ├── FirePro D700 (GCN 1.0, 2013) ← TESTED
        ├── RX 580 (GCN 4, 2017) ← TESTED
        ├── Vega FE (GCN 5, 2017) ← TESTED
        ├── RX 6600 XT / 6900 XT (RDNA 2, 2020–2021) ← TESTED
        └── RX 9070 XT (RDNA 4, 2025) ← TESTED

What This Means for the Benchmark

The Adreno 640 in this benchmark is, in a historical sense, a distant cousin of the Radeon GPUs it is compared against. Both trace their origins to ATI Technologies, but their architectures diverged completely after the 2008 sale to Qualcomm. Comparing them side-by-side in the same benchmark highlights:

  1. How far mobile GPUs have come: The Adreno 640 (a 2019 mobile GPU in a tablet SoC) can run the same Vulkan 1.1 compute + render pipeline as desktop GPUs, achieving 106 FPS — slower than a discrete desktop GPU, but functional and measurable with identical code.

  2. The power/performance trade-off: The Adreno 640 operates within a ~3W thermal envelope (tablet SoC), while the RX 9070 XT consumes ~300W. The 9070 XT is ~17× faster (1,751 vs 106 FPS), but uses ~100× more power — making the Adreno 640 significantly more performance-per-watt efficient for this workload.

  3. Driver maturity gap: Qualcomm's Windows Vulkan driver is relatively young (first WoA devices shipped 2023), while AMD's Radeon Vulkan drivers have been refined since 2016. This is reflected in the Adreno 640's higher render times and less efficient presentation path.


Appendix: Qualcomm Adreno 640 — Windows on ARM Benchmark Results

First-ever inclusion of a mobile Qualcomm GPU in this benchmark suite. The Adreno 640 runs natively on Windows 11 ARM64 via Qualcomm's WoA Vulkan/DX12/DX11 drivers — no emulation or translation layer for the GPU workload itself.

Why Include an Adreno GPU in a Desktop GPU Benchmark?

The test device is a Xiaomi Pad 5 — an Android tablet originally shipping with MIUI 12.5 based on Android 11. Through community-developed custom firmware (Project Renegade / Windows on ARM for Snapdragon 855/860 tablets), Windows 11 ARM64 was installed on this device, replacing the Android operating system entirely. The tablet boots directly into Windows 11, with Qualcomm-provided WDDM drivers exposing the Adreno 640 as a standard Windows GPU — complete with Vulkan, DirectX 12, and DirectX 11 support.

The motivation for testing this device is rooted in the ATI → AMD → Qualcomm lineage documented in the previous section. The Adreno GPU traces its origin to ATI's Imageon mobile GPU division, which AMD sold to Qualcomm in 2008. Qualcomm renamed Imageon to Adreno — an anagram of "Radeon" — and has since developed it into a fully independent architecture. While Adreno shares no silicon or ISA with modern Radeon GPUs, the historical connection means this benchmark suite now covers both branches of ATI's GPU family tree: the Radeon line (TeraScale 2 → GCN → RDNA 4) and the Adreno line that diverged in 2008.

Including the Adreno 640 also provides unique data points not available from any desktop GPU:

  1. Mobile vs desktop GPU scaling: How does a 2-5W tablet SoC GPU compare against 75–300W discrete desktop GPUs running the exact same compute + render workload?
  2. Qualcomm driver maturity: Qualcomm's Windows GPU drivers are relatively new (first WoA devices shipped 2023). How do they perform compared to AMD and NVIDIA's decade-old Windows driver stacks?
  3. ARM64 vs x64 binary translation: The same benchmark compiled as ARM64 (native) and x64 (Prism emulation) on identical hardware reveals the exact CPU-side overhead of Microsoft's binary translation layer — with GPU times serving as a constant control variable.

Test Hardware

Component Specification
Device Xiaomi Pad 5 (nabu) — Android tablet running Windows 11 ARM64 via Renegade Project / Port-Windows-11-Xiaomi-Pad-5
Previous OS Android 15 (HyperOS 2)
Current OS Windows 11 Pro ARM64 (NT 10.0.26100)
SoC Qualcomm Snapdragon 860 (SM8150-AC) — a binned Snapdragon 855+
CPU Kryo 585: 1× A77 @ 2.96 GHz (prime) + 3× A77 @ 2.42 GHz (performance) + 4× A55 @ 1.80 GHz (efficiency)
GPU Qualcomm Adreno 640, ~585 MHz boost, 384 ALUs
RAM 12 GB LPDDR4X (shared between CPU and GPU)
VRAM Shared (reported as 1 MB by driver — a Qualcomm WDDM driver reporting limitation, not actual VRAM size)
Vulkan 1.1.276 (Qualcomm proprietary driver, build 2023-10-23)
DX12 Feature Level 12_1 (driver 27.20.2060.0)
DX11 Supported (driver 27.20.2060.0)
OpenGL Not supported — Qualcomm's Windows driver does not expose desktop OpenGL. Microsoft's OpenCL/OpenGL Compatibility Pack provides only GL 3.3 via Mesa-on-DX12 translation, below this benchmark's GL 4.3 requirement. (The same hardware supports OpenGL ES 3.2 on Android, and the open-source Mesa Freedreno driver achieves full OpenGL 4.6 on Linux.)
Display 11" 2560×1600 IPS 120 Hz (benchmark runs at 1280×720)
Thermal Passive cooling only (tablet form factor, no fan)
Resolution 1280 × 720
V-Sync OFF

Cross-API Results — 1M Particles (Medium)

# API Avg FPS Compute (ms) Render (ms) Total GPU (ms) Frame Time (ms) GPU Util Bottleneck
1 DX12 116.7 3.281 4.767 8.053 8.57 94% GPU-bound
2 Vulkan 105.8 3.299 5.704 9.002 9.45 95% GPU-bound
3 DX11 83.2 4.374 4.579 11.775 12.02 98% GPU-bound

Key Observations

1. DX12 is the fastest API on Adreno 640

Unlike desktop GPUs where DX11 or Vulkan often lead, DX12 is the clear winner on this mobile GPU. This suggests Qualcomm's DX12 driver path is more optimised than their Vulkan or DX11 paths — plausible given that DX12 is the primary API for Windows on ARM gaming and application compatibility.

2. Vulkan render time is inflated

Vulkan's render time (5.704 ms) is 20% higher than DX12's (4.767 ms), despite both rendering the same workload. This likely reflects immaturity in Qualcomm's Vulkan presentation/swapchain path on Windows, similar to the semaphore wait pollution observed on desktop GPUs (Section 6) but more pronounced.

3. DX11 has the highest compute overhead

DX11 compute time (4.374 ms) is 33% higher than Vulkan/DX12 (~3.3 ms). Combined with a large total GPU time (11.775 ms), DX11 is clearly the least efficient path. The implicit driver overhead that helps DX11 on mature desktop drivers (NVIDIA, AMD) does not translate to Qualcomm's younger driver stack.

4. All APIs are GPU-bound at 95%+ utilisation

The Adreno 640 is fully saturated at 1M particles — there is no CPU overhead headroom. This contrasts sharply with desktop GPUs where most APIs are CPU-bound at this particle count (e.g., RTX 5090 at 33% GPU utilisation). The mobile GPU's lower compute throughput means it hits the GPU-bound regime much earlier.

Adreno 640 vs Desktop GPUs — Cross-Platform Comparison

GPU Best API Best FPS Compute (ms) Total GPU (ms) vs Adreno 640
RTX 5090 DX12 5,603 0.014 0.065 48× faster
RX 9070 XT DX11 1,774 0.047 0.542 15× faster
RX 6600 XT DX12 1,834 0.190 0.406 16× faster
Vega FE DX12 1,716 0.219 0.452 15× faster
RX 580 DX12 912 0.362 0.930 8× faster
GTX 970 Vulkan 719 0.434 1.098 6× faster
FirePro D700 Vulkan 555 0.589 1.473 5× faster
Radeon iGPU (2 CU) DX12 324 1.480 2.953 3× faster
HD 5770 OpenGL 188 1.794 4.818 1.6× faster
Adreno 640 DX12 117 3.281 8.053 1.0× (baseline)
WARP on 9800X3D¹ DX12 86.6 1.034 11.059 0.7× (slower)
WARP on 7600¹ DX12 62.2 1.810 15.181 0.5× (slower)
WARP on SD 860 (native)¹ DX12 4.2 30.416 187.321 0.036× (slower)
WARP on SD 860 (x64 emulated)¹ DX12 2.9 56.804 338.073 0.025× (slower)

¹ WARP is a CPU software renderer — performance depends entirely on the host CPU. The Ryzen 7 9800X3D (8-core, 96 MB 3D V-Cache) is 39% faster than the Ryzen 5 7600 (6-core, 32 MB L3), and both x86 desktop CPUs are dramatically faster than the Snapdragon 860 (1+3+4 ARM cores). On the SD 860, WARP was tested twice: the native ARM64 build (enumerated as "Microsoft Basic Render Driver") achieves 4.2 FPS, while the x64 emulated build (enumerated as "Microsoft WARP (CPU Software Renderer)", running through Prism translation) achieves only 2.9 FPS — a 45% penalty from binary translation on an already slow CPU. Even the native WARP result is 28× slower than the Adreno 640 hardware GPU on the same device, demonstrating the importance of hardware-accelerated GPU compute on mobile SoCs.

The Adreno 640 sits between the HD 5770 (a 2009 discrete desktop GPU) and the WARP software renderer in absolute performance. It outperforms the fastest WARP result (9800X3D) by 35%, confirming it is a real hardware GPU despite its mobile origins.

API Support Limitations on Windows ARM

API Adreno 640 (WoA) Desktop GPU (x64) Notes
Vulkan 1.1 (native driver) 1.3+ Qualcomm provides a native ARM64 Vulkan ICD
DX12 FL 12_1 (native driver) FL 12_1–12_2 Full native support via WDDM driver
DX11 Supported (native driver) Supported Full native support
OpenGL 3.3 max (compatibility pack) 4.6 No native desktop OpenGL; Microsoft's compatibility pack (Mesa → DX12 translation) maxes out at GL 3.3, below the 4.3 required by this benchmark
Metal N/A N/A (macOS only) —

The lack of OpenGL 4.3 support means the Adreno 640 cannot run this benchmark's OpenGL backend. This is a platform limitation, not a hardware one — the same Adreno 640 on Android supports OpenGL ES 3.2 (roughly equivalent to desktop GL 4.3 in capability), and on Linux the open-source Mesa Freedreno driver achieves full OpenGL 4.6 on Adreno 600-series hardware.


Appendix: ARM64 Native vs x64 Emulated — Performance Comparison on Adreno 640

Windows on ARM runs native ARM64 binaries at full speed, but can also execute x86/x64 applications through Microsoft's built-in binary translation layer (Prism). This section compares the same benchmark compiled as ARM64 vs x64 on identical hardware.

What is Prism (x64 Emulation on ARM)?

Windows 11 on ARM includes Prism, a binary translation layer that converts x86/x64 instructions to ARM64 at runtime. This allows unmodified x64 Windows applications to run on ARM hardware with a performance penalty. Prism translates code JIT (just-in-time), caching translated blocks for reuse.

Key characteristics:

  • CPU code is translated: all C++ application logic (particle initialisation, frame loop, API calls) runs through the translation layer
  • GPU code is NOT translated: shaders (SPIR-V, HLSL, GLSL) execute natively on the GPU regardless of the host binary's architecture
  • API calls pass through: Vulkan/DX12/DX11 driver calls from the x64 binary reach the same native ARM64 GPU driver via interop thunks

Why Compare ARM64 vs x64?

  1. RenderDoc compatibility: RenderDoc only ships as x64. An x64 build is required for RenderDoc frame capture on WoA hardware. Measuring the emulation penalty tells us whether x64 benchmark results are still meaningful for comparison.
  2. Real-world relevance: Many Windows applications remain x64-only. Understanding the GPU benchmark impact of emulation helps assess whether WoA devices can be trusted for performance-sensitive x64 workloads.
  3. Isolating CPU vs GPU overhead: Since shaders run natively regardless of binary architecture, any performance difference is purely CPU-side (API call overhead, frame loop, buffer management). This cleanly separates translation overhead from GPU execution time.

Results — ARM64 Native vs x64 Emulated (Adreno 640, 1M Particles, Medium)

Metric ARM64 Native x64 Emulated Delta
Vulkan
Avg FPS 105.8 78.6 −25.7%
Compute (ms) 3.299 3.119 −5.5%
Render (ms) 5.704 5.586 −2.1%
Total GPU (ms) 9.002 8.706 −3.3%
Frame Time (ms) 9.448 12.726 +34.7%
DX12
Avg FPS 116.7 98.3 −15.8%
Compute (ms) 3.281 3.257 −0.7%
Render (ms) 4.767 4.368 −8.4%
Total GPU (ms) 8.053 7.630 −5.3%
Frame Time (ms) 8.570 10.169 +18.7%
DX11
Avg FPS 83.2 75.2 −9.6%
Compute (ms) 4.374 4.381 +0.2%
Render (ms) 4.579 4.470 −2.4%
Total GPU (ms) 11.775 11.678 −0.8%
Frame Time (ms) 12.024 13.293 +10.6%

Analysis

GPU compute/render times are virtually identical between ARM64 and x64 builds — within ±5% across all APIs, well within run-to-run variance. This confirms the prediction: GPU shaders execute natively on the Adreno 640 regardless of the host binary's architecture. The translation layer does not affect GPU-side workload execution.

FPS is 10–26% lower on x64, with the penalty varying by API:

API FPS Penalty Frame Time Increase CPU Overhead Added
Vulkan −25.7% +3.28 ms ~3.3 ms per frame
DX12 −15.8% +1.60 ms ~1.6 ms per frame
DX11 −9.6% +1.27 ms ~1.3 ms per frame

Why Vulkan is penalised most heavily:

Vulkan has the highest per-frame CPU call count of the three APIs — explicit command buffer recording, descriptor set binding, fence management, and swapchain acquisition each require individual API calls that pass through the Prism translation layer. Each translated call adds a small overhead (~microseconds), but at 100+ calls per frame, the total accumulates to ~3.3 ms.

DX12 has slightly fewer per-frame CPU calls due to its command list model, resulting in a smaller 1.6 ms penalty. DX11's implicit driver handles most resource management internally (fewer API calls from the application), so it suffers the least translation overhead at 1.3 ms.

Why GPU times are slightly lower on x64 (counter-intuitive):

The x64 build's GPU compute and render times are marginally lower (by 1–5%) than ARM64. This is not because x64 code makes the GPU faster — it is an artefact of the higher frame time. With longer gaps between frame submissions (due to CPU translation overhead), the GPU has slightly more time to process each frame without contention, resulting in marginally cleaner timestamps. The difference is within measurement noise and should not be interpreted as a real GPU performance improvement.

Conclusions

  1. GPU benchmark data from x64 builds is valid. GPU compute and render times are unaffected by Prism translation — x64 results can be directly compared against ARM64 results for GPU performance analysis.

  2. FPS comparisons require a correction factor. x64 FPS is 10–26% lower than ARM64 native due to CPU-side translation overhead. When comparing Adreno 640 FPS against desktop GPUs (which run x64 natively), the ARM64 native results should be used as the true performance baseline.

  3. Prism overhead is workload-dependent. API-heavy workloads (Vulkan, DX12) are penalised more than API-light workloads (DX11). For GPU-bound scenarios (this benchmark at 1M particles on Adreno 640), the penalty is modest because most frame time is spent on GPU execution, not CPU API calls.

  4. RenderDoc x64 captures on WoA are viable. Since the x64 build produces identical GPU behaviour with only a CPU overhead penalty, RenderDoc frame captures from x64 builds are representative of true GPU workload behaviour — the captured GPU commands, timings, and resource state will match what the ARM64 native build would produce.