- RX 9070 XT and RX 6900 XT — Cross-API, Particle Scaling, and Headless
- Swapchain Throttling, Timestamp Pollution, and Headless Compute
- Cross-API Comparison — RTX 5090
- Cross-GPU Comparison — Vulkan
- Memory Allocation Impact — Vulkan, RTX 5090
- Software Renderer Baseline — WARP
- Legacy Discrete GPU vs Modern iGPU
- OpenGL GPU Selection — Platform Limitations
- DX11 Timestamp Query Failures
- OpenGL Compute Shader Performance on AMD GPUs
- Dual Identical GPU Behaviour (Mac Pro 2013)
- ATI, AMD, and Qualcomm Adreno — A Shared GPU Heritage
- Qualcomm Adreno 640 — Windows on ARM Benchmark Results
- ARM64 Native vs x64 Emulated — Performance Comparison
A cross-platform GPU compute and rendering microbenchmark written in C++17. It simulates millions of particles on the GPU using a compute shader, then renders them as point-sprites using a graphics pipeline — all within a single command buffer per frame.
The project implements five interchangeable graphics API backends:
| Backend | API | Shader Language | Platforms |
|---|---|---|---|
| Vulkan 1.1+ | Explicit, low-level | GLSL → SPIR-V | Windows, Linux, macOS (MoltenVK) |
| DirectX 12 | Explicit, low-level | HLSL (SM 5.1) | Windows 10+ |
| DirectX 11 | Implicit, driver-managed | HLSL (SM 5.0) | Windows 7+ |
| OpenGL 4.3 | Implicit, driver-managed | GLSL 430 | Windows, Linux |
| Metal | Explicit (Apple) | MSL | macOS |
All five backends share the same AppBase class for windowing (GLFW), particle
initialisation, timing, and benchmark result management. Each backend overrides
InitBackend(), DrawFrame(), CleanupBackend(), and WaitIdle().
Beyond the original particle simulation, the benchmark runs five interchangeable
workloads, each isolating a different axis of GPU performance and reporting a
deterministic, cross-API-comparable metric (not just FPS). Select with
--workload; every backend runs the same algorithm so numbers compare directly.
| Axis | Workload (--workload) |
What it stresses | Metric | Design |
|---|---|---|---|---|
| Bandwidth | stream (default) |
Memory subsystem | GB/s | Particle update — ~0.15 FLOP/byte, bandwidth-bound |
| Compute (achievable) | nbody |
FP32 ALU + SFU + shared memory | GFLOP/s | All-pairs N-body, tiled through shared memory (--bodies) |
| Fill / fragment | stress |
Rasteriser + fragment ALU/SFU + ROP | G-iter/s | Fullscreen fractal, fixed per-pixel iterations (--iter) |
| Compute (peak) | synthpeak |
Raw ALU throughput per precision | GFLOPS / GIOPS | Register-resident FMA loop, vkpeak-style (--precision, --iter) |
| 3D render | render3d |
Vertex transform + rasterisation + fill + depth | MQuad/s | Perspective + orbiting camera + depth test; instanced camera-facing billboard quads (--particles) |
Why both nbody and synthpeak? synthpeak measures the theoretical ceiling
(near-peak FLOPS), while nbody measures achievable compute under real data
dependencies and special-function (rsqrt) load. Together they bracket compute
performance; stream and stress add the bandwidth and fill axes; render3d adds
a real 3D graphics pipeline (the others render nothing or only a 2D pass).
Bandwidth / Compute / Fill — same workload across APIs:
| Workload | Vulkan | DX12 | DX11 | OpenGL |
|---|---|---|---|---|
| N-body (GFLOP/s, 64K bodies) | 46,681 | 36,074 | 49,094¹ | 53,798 |
| Fractal (G-iter/s, 8000 iter) | 3,681 | 3,450 | 3,445¹ | 3,521 |
| Render3D (MQuad/s, 1M billboards) | 1,861 | 1,978 | 1,655 | 1,774 |
¹ DX11 figures from windowed mode (see DX11 timestamp note below).
The Render3D pass (instanced billboards + depth + perspective) converges across APIs much like the fill/peak axes — all four are within ~20%, since the work is dominated by hardware vertex/raster/fill throughput rather than API overhead.
Synthetic peak by precision (GFLOPS / GIOPS):
| Precision | Vulkan | DX12 | OpenGL | DX11 | Notes |
|---|---|---|---|---|---|
| FP32 | 112,615 | 109,016 | 98,687 | 109,737 | ≈ 5090's ~105 TFLOPS spec |
| FP16 | 122,412 | 119,757 | 118,934 | N/A² | ≈ FP32×1.05 — consumer Blackwell has no 2× non-tensor FP16 |
| FP64 | 1,980 | 1,973 | 1,971 | 1,667 | ≈ 1/57 of FP32 (consumer 1/64 double rate) |
| INT32 | 61,508 | 56,978 | 61,557 | 63,147 | ~0.55× FP32 (IMAD ≈ half-rate) |
² DX11 cannot run true FP16: Direct3D 11 caps at Shader Model 5, whose
min16float is minimum-precision (the driver may run it at FP32), not IEEE FP16.
True 16-bit needs SM 6.2, which only the DX12 runtime exposes.
Key cross-API insight: the peak (
synthpeak) and fill (stress) axes converge tightly across APIs (within ~6%) because they are pure hardware-throughput limited — the driver/API barely matters. The achievable compute (nbody) axis shows a larger spread (~1.5×, OpenGL fastest, DX12 slowest on this NVIDIA driver) because dispatch/scheduling overhead is exposed. Running multiple axes is what separates "hardware ceiling" from "API/driver overhead."
FP16 has no single portable path, so each API uses the right tool for true 16-bit:
- Vulkan — standard path:
VK_KHR_shader_float16_int8+GL_EXT_shader_explicit_arithmetic_types_float16, packedf16vec2. ✅ - DX12 — FXC's
min16floatis only minimum-precision (often FP32). True FP16 needs Shader Model 6.2, so the FP16 kernel is precompiled to signed DXIL with the Windows SDK DXC (-enable-16bit-types) at build time and loaded at runtime; the device is checked forNative16BitShaderOps. ✅ - OpenGL — desktop GL core has no portable FP16; the FP16 kernel uses
GL_NV_gpu_shader5(float16_t/f16vec2), gated at runtime by the extension (NVIDIA). AMD would useGL_AMD_gpu_shader_half_float. ✅ (NVIDIA) - DX11 — impossible: Direct3D 11 caps at SM 5, which has no true FP16. N/A.
- Metal —
half/half2is native; written but untested here (no macOS). ✅ (code)
Likewise FP64 is unavailable on Metal — Apple GPUs have no double-precision units.
synthpeak (and the headless compute path generally) forces headless mode (no
window/swapchain). On the test driver, DX11 never resolves GPU timestamp queries
in headless mode — a pre-existing, documented behaviour (see
woa-dx11-timestamp-issue.md). The DX11 SynthPeak
kernel does execute and produce FPS, but with no GPU-time it cannot report a
GFLOPS score, so it shows as N/A. DX11 timestamps work normally in windowed
mode — which is why DX11's N-body and Fractal figures above (windowed) are present
while its headless SynthPeak is not. This is a DX11/driver limitation, not an
untested path.
All benchmarks in this report were collected on four physical systems using the GPUs and CPUs listed below.
Nine AMD GPUs spanning four architecture generations and 15 years (2009–2024), including seven discrete GPUs and two integrated GPUs:
| GPU | Architecture | Year | CU / SP | FP32 (TFLOPS) | Memory |
|---|---|---|---|---|---|
| RX 9070 XT | RDNA 4 | 2025 | 64 CU (4096 SP) | 48.70 | 16 GB GDDR6, 640 GB/s |
| RX 6900 XT | RDNA 2 | 2020 | 80 CU (5120 SP) | 23.04 | 16 GB GDDR6, 512 GB/s |
| RX 6600 XT | RDNA 2 | 2021 | 32 CU (2048 SP) | 10.60 | 8 GB GDDR6, 256 GB/s |
| Vega Frontier Edition | GCN 5 (Vega) | 2017 | 64 CU (4096 SP) | 13.11 | 16 GB HBM2, 483 GB/s |
| RX 580 | GCN 4 (Polaris 20) | 2017 | 36 CU (2304 SP) | 6.17 | 8 GB GDDR5, 256 GB/s |
| FirePro D700 | GCN 1 (Tahiti) | 2013 | 32 CU (2048 SP) | 3.48 | 6 GB GDDR5, 264 GB/s |
| HD 5770 | TeraScale 2 (Juniper) | 2009 | 10 SIMD (800 SP) | 1.36 | 1 GB GDDR5, 77 GB/s |
| Ryzen 7 9800X3D iGPU | RDNA 2 | 2024 | 2 CU (128 SP) | 0.56 | Shared DDR5, ~90 GB/s |
| Ryzen 5 7600 iGPU | RDNA 2 | 2023 | 2 CU (128 SP) | 0.56 | Shared DDR5, ~90 GB/s |
RDNA 4 uses dual-issue FP32; traditional single-issue calculation yields 24.3 TFLOPS. The two iGPUs are architecturally identical (RDNA 2, 2 CU, 2200 MHz) but sit in different CPU platforms (Zen 5 3D V-Cache vs Zen 4).
| GPU | Architecture | Year | CUDA Cores | FP32 (TFLOPS) | Memory |
|---|---|---|---|---|---|
| RTX 5090 | Blackwell | 2025 | 21,760 | 104.8 | 32 GB GDDR7, 1792 GB/s |
| GTX 970 | Maxwell 2.0 | 2014 | 1,664 | 3.92 | 4 GB GDDR5, 224 GB/s |
| GPU | Architecture | Year | ALU | FP32 (TFLOPS) | Memory |
|---|---|---|---|---|---|
| Adreno 640 | Adreno 6xx | 2019 | 768 | ~0.90 | Shared LPDDR4X, ~34 GB/s |
Tested on a Xiaomi Pad 5 (Snapdragon 860) running Windows on ARM. Supports DX11 FL 11_1, DX12 FL 12_1, and Vulkan 1.1. Microsoft WARP (software renderer) is also used as a baseline — see Section 12.
Also, All CPUs were tested with Microsoft WARP (DX11/DX12 software rasteriser) to measure CPU-side compute and rendering performance.
| CPU | Architecture | Year | Cores / Threads | Boost Clock | TDP | Platform |
|---|---|---|---|---|---|---|
| AMD Ryzen 7 9800X3D | Zen 5 + 3D V-Cache | 2024 | 8C / 16T | 5.2 GHz | 120 W | AM5, DDR5-6000 |
| AMD Ryzen 5 7600 | Zen 4 | 2023 | 6C / 12T | 5.1 GHz | 65 W | AM5, DDR5-6000 |
| Intel Xeon E5-2697 v2 | Ivy Bridge-EP | 2013 | 12C / 24T | 3.5 GHz | 130 W | LGA 2011, DDR3-1866 |
| Qualcomm Snapdragon 860 | Kryo 485 (A76/A55) | 2021 | 1P+3P+4E | 2.96 GHz | 2~7 W | LPDDR4X, WoA |
- Multi-GPU support: enumerate and select from all available GPUs
(discrete → integrated → software), with
--gpuCLI override. - GPU timestamp profiling: per-frame compute / render / total GPU timing
via
VkQueryPool(Vulkan),ID3D12QueryHeap(DX12),ID3D11Query(DX11),glQueryCounter(OpenGL). - Benchmark mode: fixed frame count or timed run, warmup period, min/max/avg statistics, CPU-bound vs GPU-bound analysis.
- Result persistence: auto-save to
~/.gpu_bench/results.json, compare, delete, CSV export. - RenderDoc integration:
VK_EXT_debug_utilslabels + In-Application API for programmatic frame capture (--capture <frame>). - Python tooling: chart generation, batch benchmark automation, markdown/HTML report export, 3DMark cross-validation.
Every frame executes the following sequence in a single command buffer (shown for the Vulkan backend; other backends follow the same logical structure):
┌──────────────────────────────────────────────────────┐
│ Command Buffer │
│ │
│ ┌─ Timestamp T0 (TOP_OF_PIPE) ───────────────────┐ │
│ │ │ │
│ │ ┌─ Particle Compute ────────────────────────┐ │ │
│ │ │ Bind compute pipeline │ │ │
│ │ │ Bind descriptor set (SSBO) │ │ │
│ │ │ Push constants (deltaTime, damping) │ │ │
│ │ │ vkCmdDispatch(N/256, 1, 1) │ │ │
│ │ └───────────────────────────────────────────┘ │ │
│ │ │ │
│ ├─ Timestamp T1 (COMPUTE_SHADER) ────────────────┤ │
│ │ │ │
│ │ ┌─ SSBO Barrier ───────────────────────────┐ │ │
│ │ │ srcStage: COMPUTE_SHADER_BIT │ │ │
│ │ │ dstStage: VERTEX_INPUT_BIT │ │ │
│ │ │ SHADER_WRITE → VERTEX_ATTRIBUTE_READ │ │ │
│ │ └──────────────────────────────────────────┘ │ │
│ │ │ │
│ ├─ Timestamp T2 (TOP_OF_PIPE) ───────────────────┤ │
│ │ │ │
│ │ ┌─ Particle Render ────────────────────────┐ │ │
│ │ │ Begin render pass (clear to dark blue) │ │ │
│ │ │ Bind graphics pipeline │ │ │
│ │ │ Bind vertex buffer (same SSBO) │ │ │
│ │ │ vkCmdDraw(particleCount, 1, 0, 0) │ │ │
│ │ │ End render pass │ │ │
│ │ └──────────────────────────────────────────┘ │ │
│ │ │ │
│ └─ Timestamp T3 (COLOR_ATTACHMENT_OUTPUT) ────────┘ │
│ │
└──────────────────────────────────────────────────────┘
The compute shader (shaders/compute.comp) performs Euler integration on each
particle:
layout(local_size_x = 256) in;
layout(set = 0, binding = 0, std430) buffer ParticleBuffer {
Particle particles[]; // vec4 position + vec4 velocity = 32 bytes
};
layout(push_constant) uniform ComputeParams {
float deltaTime;
float bounds;
};
void main() {
uint i = gl_GlobalInvocationID.x;
particles[i].position.xyz += particles[i].velocity.xyz * deltaTime;
if (particles[i].position.x > bounds)
particles[i].position.x = -bounds;
}- Workgroup size: 256 threads — balances occupancy across AMD (wavefront 64) and NVIDIA (warp 32) architectures.
- SSBO layout:
std430guarantees C++-compatible packing (no padding). - Push constants: avoid descriptor set updates every frame; only 8 bytes pushed per dispatch.
For 16M particles: 16,777,216 / 256 = 65,536 workgroups dispatched.
The render pipeline draws all particles as POINT_LIST using the same SSBO
as a vertex buffer:
| Stage | Configuration |
|---|---|
| Vertex Input | Binding 0, stride 32 bytes: position (R32G32_SFLOAT, offset 0), colour (R32G32B32A32_SFLOAT, offset 16) |
| Vertex Shader | Maps particle speed to a blue → red colour gradient, sets gl_PointSize = 2.0 |
| Rasteriser | Point topology, no culling |
| Fragment Shader | Passes interpolated colour to output |
| Colour Blend | Additive blending (SRC_ALPHA + ONE) — overlapping particles create bright clusters |
| Render Pass | Single subpass, one colour attachment (swapchain image, B8G8R8A8_SRGB) |
| Mode | Vulkan Flags | When Used |
|---|---|---|
| Device-local (default) | DEVICE_LOCAL + staging buffer copy |
Discrete GPU — particle data lives in VRAM |
Host-visible (--host-memory) |
HOST_VISIBLE | HOST_COHERENT |
Integrated GPU or debugging — CPU-mappable, no staging copy |
On discrete GPUs, the staging buffer is created, filled with initial particle
data, copied via vkCmdCopyBuffer, then destroyed. All subsequent compute
and render operations access only the device-local buffer.
The single VkBufferMemoryBarrier between compute and render ensures:
- All compute shader writes to particle positions are visible before the vertex shader reads them.
- No additional barriers are needed because the entire frame is recorded into one command buffer on one queue.
- On integrated GPUs with unified memory, the barrier is essentially a no-op (no cache flush between separate memory domains).
Each backend implements GPU timestamp queries using a ring buffer of
kTimestampSlotCount (8) frame slots to avoid blocking on in-flight frames:
| API | Write | Read | Sync | Clock Frequency |
|---|---|---|---|---|
| Vulkan | vkCmdWriteTimestamp |
vkGetQueryPoolResults |
Fence wait from previous frame | timestampPeriod from device properties |
| DX12 | ID3D12GraphicsCommandList::EndQuery |
ResolveQueryData + readback buffer |
Fence signal/wait | ID3D12CommandQueue::GetTimestampFrequency |
| DX11 | ID3D11DeviceContext::End(query) |
GetData with retry loop |
Disjoint query (S_FALSE → retry) |
D3D11_QUERY_DATA_TIMESTAMP_DISJOINT.Frequency |
| OpenGL | glQueryCounter(GL_TIMESTAMP) |
glGetQueryObjectui64v |
GL_QUERY_RESULT_AVAILABLE poll |
Fixed 1 ns resolution |
Four timestamps per frame yield three intervals: compute time (T1−T0), render time (T3−T2), and total GPU time (T3−T0).
Industry-standard GPU profiling of the Vulkan particle compute + render pipeline using RenderDoc — the most widely used cross-vendor, cross-API GPU frame debugger.
RenderDoc is a free, open-source (MIT licence) GPU frame debugger created by Baldur Karlsson. It intercepts a single frame's worth of graphics API calls, allowing post-mortem inspection of every resource, pipeline state, and GPU operation. It is one of the tools referenced by AMD's JD under "Use of industry-standard profiling and debug tools".
| Tool | Vendor | Platform |
|---|---|---|
| RenderDoc | Open-source | Vulkan, DX11, DX12, OpenGL — all GPUs |
| PIX | Microsoft | DX12 (Windows only) |
| Radeon GPU Profiler (RGP) | AMD | Vulkan, DX12 (AMD GPUs only) |
| Nsight Graphics | NVIDIA | Vulkan, DX, OpenGL (NVIDIA GPUs only) |
| Xcode GPU Debugger | Apple | Metal (macOS / iOS only) |
The Vulkan backend integrates two layers of RenderDoc support:
VK_EXT_debug_utils— debug labels and object names baked into the command buffer, providing readable annotations inside RenderDoc's event browser.- RenderDoc In-Application API (
renderdoc_app.h) — runtime detection of RenderDoc, enabling F12 manual capture and--capture <frame>CLI auto-capture from within the application.
A single frame records the following sequence into one VkCommandBuffer:
| # | Event | Debug Label | Description |
|---|---|---|---|
| 1 | vkCmdResetQueryPool |
— | Reset timestamp query slots for this frame |
| 2 | vkCmdWriteTimestamp (TOP_OF_PIPE) |
— | T0: frame start |
| 3 | vkCmdBindPipeline (COMPUTE) |
Particle Compute (green) | Bind compute pipeline |
| 4 | vkCmdBindDescriptorSets |
Bind SSBO descriptor (set 0, binding 0) | |
| 5 | vkCmdPushConstants |
Push deltaTime (float) + damping (float) |
|
| 6 | vkCmdDispatch(N/256, 1, 1) |
Dispatch compute workgroups | |
| 7 | vkCmdWriteTimestamp (COMPUTE_SHADER) |
— | T1: compute end |
| 8 | vkCmdPipelineBarrier |
SSBO Barrier (yellow) | SHADER_WRITE → VERTEX_ATTRIBUTE_READ |
| 9 | vkCmdWriteTimestamp (TOP_OF_PIPE) |
— | T2: render start |
| 10 | vkCmdBeginRenderPass |
Particle Render (blue) | Clear to (0.04, 0.08, 0.14, 1.0) |
| 11 | vkCmdBindPipeline (GRAPHICS) |
Bind graphics pipeline | |
| 12 | vkCmdBindVertexBuffers |
Bind particle SSBO as VBO | |
| 13 | vkCmdDraw(particleCount, 1, 0, 0) |
Draw all particles as POINT_LIST |
|
| 14 | vkCmdEndRenderPass |
Finish render pass | |
| 15 | vkCmdWriteTimestamp (COLOR_OUTPUT) |
— | T3: render end |
| Object | VK_EXT_debug_utils Name |
Purpose |
|---|---|---|
particleBuffer_ |
Particle SSBO |
Interleaved position + velocity + colour buffer, used as both compute SSBO and vertex VBO |
computePipeline_ |
Compute Pipeline |
Particle physics update (gravity, damping, boundary reflection) |
graphicsPipeline_ |
Graphics Pipeline |
Point-sprite rendering with additive blending |
renderPass_ |
Main Render Pass |
Single subpass, colour-only attachment |
1. SSBO Data (Particle Buffer)
- Select the
vkCmdDispatchevent → Pipeline State → Compute Shader → Descriptor Set 0 → Binding 0. - View buffer contents with custom struct format:
float2 pos; float2 vel; float4 col; - Compare particle positions before and after the dispatch using
RenderDoc's timeline — values should change by
velocity × deltaTime. - Verify no NaN or out-of-bounds values.
2. Pipeline State
| Stage | Key Settings |
|---|---|
| Compute | local_size = (256, 1, 1), 1 SSBO, 2 push constants |
| Vertex Input | Binding 0 stride = 32 bytes; attr 0 = position (R32G32_SFLOAT, offset 0), attr 1 = colour (R32G32B32A32_SFLOAT, offset 16) |
| Rasteriser | POINT_LIST topology, point size = 2.0 (set in vertex shader) |
| Blend | Additive: srcColourBlendFactor = SRC_ALPHA, dstColourBlendFactor = ONE |
3. Barrier Correctness
The single vkCmdPipelineBarrier between compute and render ensures:
- Source:
VK_PIPELINE_STAGE_COMPUTE_SHADER_BIT/VK_ACCESS_SHADER_WRITE_BIT - Destination:
VK_PIPELINE_STAGE_VERTEX_INPUT_BIT/VK_ACCESS_VERTEX_ATTRIBUTE_READ_BIT - Scope: single buffer (
Particle SSBO),VK_WHOLE_SIZE.
No additional implicit barriers should be inserted by the driver. If any appear in RenderDoc's event list, they indicate suboptimal synchronisation.
4. GPU Timing Cross-Validation
Compare RenderDoc's built-in per-event timing against the application's own
vkCmdWriteTimestamp results:
| Metric | App Timestamps (ms) | Notes |
|---|---|---|
| Compute dispatch | 0.033 | vkCmdWriteTimestamp T0 → T1 |
| Render pass | 0.408 | vkCmdWriteTimestamp T2 → T3 (includes swapchain semaphore wait — see Section 6) |
| Total GPU time | 0.446 | T0 → T3 (RX 9070 XT, 1M particles) |
The Chrome JSON from renderdoccmd convert records CPU-side API call
durations (nanosecond resolution). For GPU-side per-event timing, open the
.rdc capture in RenderDoc GUI → Window → Performance Counter Viewer.
Deviation between app timestamps and RenderDoc GPU counters should be < 5 %.
Larger discrepancies may indicate RenderDoc interception overhead or
single-frame vs multi-frame averaging.
5. Potential Optimisations (identified via RenderDoc)
- Vulkan 1.3 barrier upgrade: Replace
VERTEX_INPUT_BITwith the more preciseVK_PIPELINE_STAGE_2_VERTEX_ATTRIBUTE_INPUT_BITto reduce stall scope. - Indirect dispatch: Replace hardcoded
vkCmdDispatch(N/256, 1, 1)withvkCmdDispatchIndirectto allow GPU-driven workload sizing. - Dynamic point size: Move the hardcoded
gl_PointSize = 2.0to a push constant for runtime adjustment without pipeline recreation.
Build first with scripts\build-windows.ps1 (or CMake). Prefer the preset CLI
path below; build\Release\... remains valid only if you configured -B build.
# Option A — GUI: launch from RenderDoc, press F12 during rendering
# Executable: out\build\windows-x64-release\Release\gpu_benchmark.exe
# (or gui\x64\Release\gpu_benchmark.exe after build-windows.ps1)
# Working dir: same folder as the exe (shaders are adjacent)
# Option B — CLI (Windows)
& "C:\Program Files\RenderDoc\renderdoccmd.exe" capture `
.\out\build\windows-x64-release\Release\gpu_benchmark.exe --backend vulkan --benchmark 200
# Option C — Auto-capture frame 50 (must be launched via RenderDoc)
.\out\build\windows-x64-release\Release\gpu_benchmark.exe --backend vulkan --benchmark 200 --capture 50# Linux
renderdoccmd capture ./build/gpu_benchmark --backend vulkan --benchmark 200The Full Analysis workflow (menu option 5/6) automatically captures one frame per
API backend via the RenderDoc In-Application API, converts each .rdc to Chrome
JSON using renderdoccmd convert, and runs rdoc_analyse.py to produce a
structural comparison across all four Windows backends.
Per-Frame Event Count Comparison (RX 9070 XT, 1M particles, Medium):
| API | Total Events | Frame Events | Dispatches | Draw Calls | Barriers |
|---|---|---|---|---|---|
| Vulkan 1.2 | 77 | 28 | 1 | 1 | 1 |
| DirectX 12 | 74 | 30 | 1 | 1 | 5 |
| DirectX 11 | 86 | 26 | 1 | 1 | 0 |
| OpenGL 4.3 | 90 | 17 | 1 | 1 | 1 |
-
DX12 requires 5 resource barriers to accomplish what Vulkan and OpenGL each handle with a single barrier. This reflects DX12's finer-grained resource state tracking — each buffer/texture transition (e.g.
UNORDERED_ACCESS → VERTEX_AND_CONSTANT_BUFFER,RENDER_TARGET → PRESENT) is an explicit barrier. Vulkan batches the same transitions into onevkCmdPipelineBarriercall with multiple memory barriers. -
DX11 has zero explicit barriers. The DX11 driver silently inserts all necessary synchronisation on behalf of the application. This is the key trade-off of implicit APIs: simpler code at the cost of opaque scheduling decisions that profiling tools like RenderDoc cannot surface.
-
OpenGL's frame has the fewest events (17 frame events) because the OpenGL driver consolidates many state changes into fewer internal calls. The single
glMemoryBarrier(GL_VERTEX_ATTRIB_ARRAY_BARRIER_BIT)maps directly to the same compute → vertex synchronisation as Vulkan'svkCmdPipelineBarrier. -
Vulkan debug labels are visible in captures — 6
vkCmdBeginDebugUtilsLabelEXT/vkCmdEndDebugUtilsLabelEXTcalls (3 pairs: "Particle Compute", "SSBO Barrier", "Particle Render") and 4vkCmdWriteTimestampcalls confirm the profiling instrumentation is correctly captured. -
DX12 has the most frame events (30) despite having fewer total events. This is because explicit APIs expose more per-frame work: command allocator reset, command list recording, root signature binding, and explicit resource transitions that implicit APIs hide from the application.
| Vulkan | DX12 | DX11 | OpenGL | |
|---|---|---|---|---|
| Barrier model | vkCmdPipelineBarrier — stage + access masks |
ResourceBarrier — state transitions |
Implicit (driver-managed) | glMemoryBarrier — bit flags |
| Barriers per frame | 1 | 5 | 0 | 1 |
| Who manages sync? | Application | Application | Driver | Application (coarse) |
| Profiler visibility | Full | Full | Hidden | Partial |
The full per-event command sequences for all APIs are in
docs/rdoc_comparison.md, auto-generated byscripts/rdoc_analyse.py.
See
docs/renderdoc-analysis.mdfor the detailed Vulkan analysis template anddocs/renderdoc-capture-guide.mdfor step-by-step capture instructions.
Automated data analysis scripts in the scripts/ directory.
| Script | Purpose |
|---|---|
plot_results.py |
Read results.json and generate 4 charts: FPS by GPU × API, GPU time breakdown, CPU overhead, particle-count scaling |
batch_benchmark.py |
Iterate over all GPU × API × particle-count combinations, invoke the benchmark executable, and collect results automatically |
export_report.py |
Export results as markdown tables or a standalone sortable HTML report with dark theme |
compare_3dmark.py |
Cross-validate against 3DMark: normalised bar chart, R² correlation scatter plot, auto-import from .3dmark-result files |
All scripts read from ~/.gpu_bench/results.json (the application's auto-saved
benchmark results). Charts use a dark colour scheme with API-specific colours
(Vulkan = red, DX12 = blue, DX11 = green, OpenGL = orange, Metal = purple).
AMD's first RDNA 4 discrete GPU, tested across three scenarios: standard 1M windowed, maximum 16M windowed, and headless compute. The 9070 XT provides a unique lens into swapchain throttling behaviour due to its very fast compute throughput relative to presentation overhead.
| Component | Specification |
|---|---|
| CPU | AMD Ryzen 5 7600 6-Core Processor |
| GPU | AMD Radeon RX 9070 XT (RDNA 4, 64 CU, 16 GB GDDR6, 640 GB/s) |
| Driver | AMD Adrenalin 26.3.1 (LLPC) |
| OS | Windows 11 (NT 10.0.26200) |
| Resolution | 1280 × 720 |
| V-Sync | OFF |
| Memory Mode | Device-local |
RX 9070 XT:
| # | API | Avg FPS | Compute (ms) | Render (ms) | Total GPU (ms) | GPU Util | Bottleneck |
|---|---|---|---|---|---|---|---|
| 1 | DX11 | 1,773.7 | 0.047 | 0.451 | 0.542 | 100% | GPU-bound |
| 2 | Vulkan | 1,750.6 | 0.033 | 0.408 | 0.446 | 80% | Balanced |
| 3 | DX12 | 1,608.7 | 0.034 | 0.399 | 0.434 | 70% | Balanced |
| 4 | OpenGL | 253.4 | 2.612 | 0.792 | 3.658 | 90% | GPU-bound |
RX 6900 XT:
| # | API | Avg FPS | Compute (ms) | Render (ms) | Total GPU (ms) | GPU Util | Bottleneck |
|---|---|---|---|---|---|---|---|
| 1 | DX11 | 4,067.9 | 0.047 | 0.140 | 0.226 | 90% | GPU-bound |
| 2 | DX12 | 3,518.0 | 0.049 | 0.112 | 0.162 | 60% | Balanced |
| 3 | Vulkan | 2,884.6 | 0.063 | 0.133 | 0.196 | 60% | Balanced |
| 4 | OpenGL | 228.6 | 2.742 | 1.151 | 4.066 | 90% | GPU-bound |
Key observations:
- At 1M particles, the 6900 XT achieves higher FPS (4,068 DX11 vs 1,774 on 9070 XT). This is again a CU-count advantage in this presentation-limited scenario — more CUs finish the trivial compute+render faster, leaving more frame budget for swapchain overhead.
- DX11 and Vulkan are nearly tied on the 9070 XT at ~1750 FPS. Unlike on RTX 5090 where DX11 dominates (8955 FPS vs 3611 Vulkan), the gap is much smaller on AMD — the AMD DX11 driver does not have the same micro-optimisation advantage as NVIDIA's.
- DX11 leads on both GPUs — on the 6900 XT (4,068 FPS), DX11 is 16% faster than DX12 (3,518 FPS) and 41% faster than Vulkan (2,885 FPS). At these extreme frame rates, DX11's lower per-frame CPU overhead dominates.
- OpenGL is severely penalised on both GPUs — the AMD OpenGL compute overhead issue (Section 16) causes ~2.6–2.8 ms compute time vs 0.03–0.05 ms on Vulkan, a ~70× penalty.
- All APIs show inflated render times (0.1–0.8 ms) due to swapchain semaphore wait pollution (detailed in Section 6). The actual render work is ~0.04 ms.
RX 9070 XT:
| # | API | Avg FPS | Compute (ms) | Render (ms) | Total GPU (ms) | GPU Util | Bottleneck |
|---|---|---|---|---|---|---|---|
| 1 | Vulkan | 110.9 | 1.963 | 6.440 | 8.418 | 90% | GPU-bound |
| 2 | DX12 | 105.7 | 1.940 | 6.126 | 8.069 | 90% | GPU-bound |
| 3 | DX11 | 93.5 | 1.908 | 6.152 | 9.904 | 90% | GPU-bound |
| 4 | OpenGL | 15.7 | 47.959 | 10.073 | 58.473 | 90% | GPU-bound |
RX 6900 XT:
| # | API | Avg FPS | Compute (ms) | Render (ms) | Total GPU (ms) | GPU Util | Bottleneck |
|---|---|---|---|---|---|---|---|
| 1 | Vulkan | 218.4 | 2.145 | 2.079 | 4.224 | 90% | GPU-bound |
| 2 | DX12 | 205.4 | 2.113 | 1.851 | 4.028 | 90% | GPU-bound |
| 3 | DX11 | 133.9 | 2.121 | 1.966 | 6.664 | 90% | GPU-bound |
| 4 | OpenGL | 14.3 | 45.609 | 19.009 | 64.759 | 90% | GPU-bound |
Side-by-side 16M comparison (9070 XT vs 6900 XT):
| API | 9070 XT FPS | 6900 XT FPS | 9070 XT Compute | 6900 XT Compute | 9070 XT Render | 6900 XT Render |
|---|---|---|---|---|---|---|
| Vulkan | 110.9 | 218.4 | 1.963 ms | 2.145 ms | 6.440 ms | 2.079 ms |
| DX12 | 105.7 | 205.4 | 1.940 ms | 2.113 ms | 6.126 ms | 1.851 ms |
| DX11 | 93.5 | 133.9 | 1.908 ms | 2.121 ms | 6.152 ms | 1.966 ms |
| OpenGL | 15.7 | 14.3 | 47.959 ms | 45.609 ms | 10.073 ms | 19.009 ms |
An interesting result: the RX 6900 XT outperforms the newer RX 9070 XT at 16M particles (218 vs 111 FPS on Vulkan), despite the 9070 XT scoring 1.4× higher in Time Spy and 1.7× higher in Steel Nomad — and being faster in both pure compute (headless, Section 5c) and 1M windowed. The compute times are nearly identical (~2.0 ms); the difference is entirely in the render pass: 6.4 ms on the 9070 XT vs 2.1 ms on the 6900 XT.
This is a workload-specific result that highlights what this microbenchmark actually measures. Rendering 16M individual point primitives is an extreme stress test for primitive throughput — the ability to set up and rasterise millions of tiny primitives per frame. The 6900 XT has 80 CUs with 1.25× more rasterisation hardware than the 9070 XT's 64 CUs, but the render time gap (3.1×) is far larger than the CU ratio alone would suggest.
Infinity Cache plays a significant role here. AMD has progressively reduced Infinity Cache capacity across RDNA generations — 128 MB (RDNA 2) → 96 MB (RDNA 3, e.g. RX 7900 XTX) → 64 MB (RDNA 4) — trading raw capacity for improved per-MB efficiency. However, in this 16M-particle scenario, the 6900 XT's 2× larger cache (128 MB vs 64 MB) is a clear advantage: 16M point primitives generate heavy vertex fetch and rasterisation traffic, and the larger cache keeps more of this data on-die, reducing round-trips to VRAM. The 9070 XT compensates with higher memory bandwidth (640 vs 512 GB/s), but bandwidth cannot fully offset the latency penalty when the working set exceeds the cache. This mirrors a well-known pattern in CPUs — AMD's Ryzen 9800X3D with 3D V-Cache (96 MB L3) dramatically outperforms the standard 9700X (32 MB L3) in cache-sensitive gaming workloads, despite identical core counts and clocks. Whether on a GPU or CPU, when the working set fits in a larger cache, the raw bandwidth of the smaller-cache part cannot compensate for the hit-rate advantage.
Beyond cache, the 6900 XT also benefits from a more mature RDNA 2 driver pipeline for point primitive rendering. Standard game rendering uses far fewer, larger triangles with complex shading — a scenario where the 9070 XT's architectural improvements (higher clocks, ray tracing hardware, improved schedulers) deliver the performance uplift reflected in 3DMark. The lesson is that no single benchmark captures all aspects of GPU performance; this microbenchmark specifically targets compute + primitive throughput, which produces a different ranking than rasterisation-focused benchmarks with complex geometry.
16× particle scaling analysis (1M → 16M) — RX 9070 XT:
| API | 1M Compute | 16M Compute | Scaling (expected 16×) | 1M FPS | 16M FPS | FPS Ratio |
|---|---|---|---|---|---|---|
| Vulkan | 0.033 | 1.963 | 59.5× | 1,750.6 | 110.9 | 15.8× |
| DX12 | 0.034 | 1.940 | 57.1× | 1,608.7 | 105.7 | 15.2× |
| DX11 | 0.047 | 1.908 | 40.6× | 1,773.7 | 93.5 | 19.0× |
| OpenGL | 2.612 | 47.959 | 18.4× | 253.4 | 15.7 | 16.1× |
- Compute time scales super-linearly (57–60× for 16× particles on Vulkan/DX12). This is expected: at 1M particles the GPU is underutilised and the per-dispatch overhead dominates; at 16M particles the ALUs and memory bandwidth are fully saturated.
- FPS scales roughly linearly (~16× reduction) because at 16M particles all APIs are GPU-bound — CPU overhead is negligible relative to GPU execution time.
- DX11's compute–render gap widens: total GPU (9.904 ms) exceeds compute + render sum (1.908 + 6.152 = 8.060 ms) by 1.844 ms, suggesting pipeline synchronisation overhead similar to what was observed on the GTX 970 (Section 16).
- OpenGL's compute time explodes to 47.959 ms — the AMD OpenGL dispatch overhead scales worse than linearly with particle count, making OpenGL completely impractical for high particle counts on AMD.
- Render time dominates at 16M: ~6 ms for Vulkan/DX12/DX11 vs ~0.4 ms at 1M. This is real render work (16M point sprites), not semaphore wait — at 16M particles the GPU is genuinely busy rendering, unlike at 1M where the semaphore wait dominated.
RX 9070 XT:
| # | API | Avg FPS | Compute (ms) | Render (ms) | Total GPU (ms) | GPU Util | Bottleneck |
|---|---|---|---|---|---|---|---|
| 1 | DX12 | 21,354.0 | 0.034 | 0.0 | 0.035 | 70% | Balanced |
| 2 | Vulkan | 21,259.6 | 0.034 | 0.0 | 0.034 | 70% | Balanced |
| 3 | OpenGL | 20,298.3 | 0.034 | 0.0 | 0.034 | 70% | Balanced |
| 4 | DX11 | 16,937.8 | 0.034 | 0.0 | 0.034 | 60% | Balanced |
RX 6900 XT:
| # | API | Avg FPS | Compute (ms) | Render (ms) | Total GPU (ms) | GPU Util | Bottleneck |
|---|---|---|---|---|---|---|---|
| 1 | DX12 | 15,949.6 | 0.048 | 0.0 | 0.049 | 70% | Balanced |
| 2 | Vulkan | 15,241.5 | 0.055 | 0.0 | 0.055 | 70% | Balanced |
| 3 | DX11 | 12,274.7 | 0.047 | 0.0 | 0.052 | 60% | Balanced |
| 4 | OpenGL | 341.2 | 2.750 | 0.0 | 2.750 | 90% | GPU-bound |
Headless cross-GPU comparison (1M particles):
| API | 9070 XT FPS | 6900 XT FPS | 9070 XT / 6900 XT |
|---|---|---|---|
| Vulkan | 21,260 | 15,242 | 1.39× |
| DX12 | 21,354 | 15,950 | 1.34× |
| DX11 | 16,938 | 12,275 | 1.38× |
| OpenGL | 20,298 | 341 | 59.5× |
In headless mode, the 9070 XT is consistently ~1.37× faster than the 6900 XT in pure compute — this is the true generational improvement, free of presentation throttling and render pipeline differences. The 9070 XT achieves this with 64 CUs vs 80 CUs (80% of the CU count), meaning its per-CU compute efficiency is ~1.7× higher than RDNA 2. The 6900 XT's OpenGL headless result (341 FPS vs 15K+ on other APIs) confirms the AMD OpenGL compute overhead persists even without a window.
Why Infinity Cache no longer matters here. In Section 5b, the 6900 XT's 128 MB Infinity Cache was identified as a key factor in its 16M-particle render advantage over the 9070 XT (64 MB). In headless mode, that advantage disappears entirely — the 9070 XT wins by 1.37×. The reason is a fundamental difference in memory access patterns: the compute shader performs a streaming update (each particle reads its own position/velocity, updates, and writes back), where data is touched once and not reused — cache size is irrelevant, and raw bandwidth (640 vs 512 GB/s) and ALU throughput determine performance. By contrast, the 16M-particle render pass involves vertex fetch, primitive assembly, rasterisation, and depth testing across the same memory regions, creating repeated, overlapping accesses where a larger cache dramatically improves hit rates. This confirms that Infinity Cache is a render-path advantage in cache-sensitive workloads, not a universal compute advantage. Paradoxically, headless compute results are a better proxy for traditional gaming rasterisation performance at this GPU tier than the 16M-particle windowed render test. Real game rendering uses thousands of large, shaded triangles — a workload dominated by ALU throughput, bandwidth, and clock speed rather than raw primitive count. The headless 1.37× advantage for the 9070 XT aligns closely with its 1.4× Time Spy and 1.7× Steel Nomad leads, while the 16M-particle render test — with its extreme small-primitive throughput stress — is an outlier that specifically punishes smaller caches and fewer fixed-function rasterisation units. In other words, the windowed 16M test tells us something real about the hardware, but it is the headless result that better predicts how these GPUs rank in the workloads most users care about.
RX 9070 XT — Windowed vs Headless comparison:
| API | Windowed FPS | Headless FPS | Speedup | Windowed Compute | Headless Compute |
|---|---|---|---|---|---|
| Vulkan | 1,750.6 | 21,259.6 | 12.1× | 0.033 ms | 0.034 ms |
| DX12 | 1,608.7 | 21,354.0 | 13.3× | 0.034 ms | 0.034 ms |
| DX11 | 1,773.7 | 16,937.8 | 9.5× | 0.047 ms | 0.034 ms |
| OpenGL | 253.4 | 20,298.3 | 80.1× | 2.612 ms | 0.034 ms |
- All four APIs converge to identical compute time (0.034 ms) on the 9070 XT in headless mode, proving the GPU compute hardware is equivalent regardless of API.
- OpenGL's 80× speedup is the most dramatic — headless mode bypasses the AMD OpenGL compute dispatch overhead entirely, as the dispatch path is simpler without a rendering context.
- DX11 compute drops from 0.047 to 0.034 ms in headless, suggesting the DX11 driver's implicit state management adds ~0.013 ms overhead even to compute dispatch when a swapchain is present.
- The remaining FPS differences (21K vs 17K for DX11) reflect pure CPU-side overhead differences between APIs.
What is "Frames-in-Flight"? In modern graphics APIs, the CPU does not wait for the GPU to finish one frame before starting the next. Instead, the CPU can prepare N frames ahead while the GPU is still rendering earlier ones — these N in-progress frames are called "frames-in-flight" (also known as "buffered frames" or "frame overlap"). This is essentially a render queue — the number of frames simultaneously in-progress in the CPU–GPU pipeline, analogous to the "work-in-progress" slots on a factory assembly line. With 2 flights, the CPU can be building frame N+1 while the GPU renders frame N; with 3 flights, the CPU can be up to 2 frames ahead. A deeper queue improves throughput (the pipeline is less likely to stall), but increases input latency (the displayed frame was prepared further in the past). This is the same concept exposed by NVIDIA's "Maximum Pre-Rendered Frames" setting (now called "Low Latency Mode") and targeted by latency-reduction technologies like NVIDIA Reflex and AMD Anti-Lag, which dynamically shorten the render queue to minimise input-to-display delay at the cost of some throughput. This test compares 2 vs 3 frames-in-flight to measure the throughput impact on each API.
Relationship with V-Sync. Frames-in-flight and V-Sync are related but distinct concepts. V-Sync locks presentation to the display refresh rate to prevent tearing; frames-in-flight controls how far the CPU can work ahead of the GPU regardless of V-Sync. The two interact most visibly when V-Sync is ON: on a 60 Hz display, 2 flights means the CPU leads by up to 1 frame (~17 ms extra latency), while 3 flights ("triple buffering") allows up to 2 frames ahead (~33 ms extra latency). With V-Sync OFF — the condition used in all tests in this report — flights purely affect CPU–GPU pipeline overlap without any refresh-rate constraint, which is why DX12 can gain +22% throughput simply by moving from 2 to 3 flights: the CPU submits work earlier instead of stalling on a busy swapchain image.
| API | Flights=2 FPS | Flights=3 FPS | Change | Flights=2 Render | Flights=3 Render |
|---|---|---|---|---|---|
| Vulkan | 1,750.6 | 1,736.9 | −0.8% | 0.408 ms | 0.409 ms |
| DX12 | 1,608.7 | 1,964.5 | +22.1% | 0.399 ms | 0.400 ms |
| DX11 | 1,773.7 | 1,956.3 | +10.3% | 0.451 ms | 0.399 ms |
| OpenGL | 253.4 | 256.1 | +1.1% | 0.792 ms | 0.781 ms |
- DX12 benefits most from an extra frame-in-flight (+22%), suggesting its command pipeline can overlap more work with 3 buffers.
- Vulkan shows no improvement — its presentation engine already manages buffering efficiently at 2 frames.
- Render times remain unchanged across both flight counts, confirming that swapchain semaphore wait pollution is not reduced by adding more swapchain images (as discussed in Section 6d).
- Why this benchmark defaults to 2 flights. For GPU benchmarking, fewer frames-in-flight is generally preferable: it minimises CPU-side pipeline overlap so that the measured frame time more closely reflects actual GPU execution cost. With 3 flights, the CPU can "hide" stalls by working further ahead, which inflates FPS without the GPU doing any more work per unit time — this is a CPU-side throughput optimisation, not a GPU performance improvement. The DX12 +22% gain from 3 flights demonstrates exactly this effect: the GPU render time is unchanged (0.399 → 0.400 ms), meaning the GPU is doing identical work; the extra FPS comes entirely from the CPU submitting commands more efficiently. For a benchmark designed to measure GPU compute and render performance, 2 flights provides a cleaner signal with less CPU-side noise. This section tests 3 flights to quantify the presentation overhead, not to suggest it as a better default.
| GPU | Architecture | Compute (ms) | Render (ms) | Total GPU (ms) | FPS |
|---|---|---|---|---|---|
| RTX 5090 | Blackwell (170 SM) | 0.025 | 0.064 | 0.090 | 3,611 |
| RX 9070 XT | RDNA 4 (64 CU) | 0.033 | 0.408 | 0.446 | 1,751 |
| RX 6900 XT | RDNA 2 (80 CU) | 0.063 | 0.134 | 0.197 | 2,866 |
| RX 6600 XT | RDNA 2 (32 CU) | 0.270 | 0.379 | 0.649 | 1,239 |
| Vega FE | GCN 5 (64 CU) | 0.219 | 0.275 | 0.495 | 1,370 |
| RX 580 | GCN 4 (36 CU) | 0.362 | 0.702 | 1.070 | 783 |
Compute performance ranking (lower is better):
| GPU | Compute (ms) | vs RX 9070 XT | Per-CU Efficiency vs 9070 XT |
|---|---|---|---|
| RTX 5090 | 0.025 | 0.76× | N/A (different arch) |
| RX 9070 XT | 0.033 | 1.00× | 1.00× |
| RX 6900 XT | 0.063 | 1.91× | 0.42× (80 CU → per-CU: 1.91 × 64/80 = 1.53× slower) |
| Vega FE | 0.219 | 6.64× | 0.15× (64 CU → per-CU: 6.64× slower) |
| RX 6600 XT | 0.270 | 8.18× | 0.24× (32 CU → per-CU: 8.18 × 64/32 = 4.1× slower) |
| RX 580 | 0.362 | 10.97× | 0.16× (36 CU → per-CU: 10.97 × 64/36 = 6.2× slower) |
The 9070 XT is 8.2× faster overall than the 6600 XT in compute. Since the 9070 XT has 2× the CU count (64 vs 32), the per-CU efficiency improvement is ~4.1×, which is still a massive generational leap from RDNA 2 to RDNA 4:
- Higher clock speed (2,970 MHz vs 2,589 MHz) accounts for ~1.15×
- 2× the CU count accounts for another ~2×
- The remaining ~1.8× improvement comes from architectural changes: improved compute scheduler, better cache hierarchy, wider memory interface utilisation (640 vs 256 GB/s), and mature RDNA 4 driver code generation
The 9070 XT also outperforms the 80-CU RX 6900 XT in compute (0.033 vs 0.063 ms) despite having 80% of its CU count, making it the fastest AMD compute GPU tested in this benchmark.
During windowed benchmark runs, the RX 9070 XT exhibited an unexpected anomaly: despite computing 2× faster than the RX 6900 XT, its reported render time was higher, resulting in lower-than-expected total GPU time efficiency.
| GPU | Compute (ms) | Render (ms) | Total GPU (ms) | FPS |
|---|---|---|---|---|
| RTX 5090 | 0.025 | 0.064 | 0.090 | 3,611 |
| RX 6900 XT | 0.063 | 0.134 | 0.197 | 2,866 |
| RX 9070 XT | 0.033 | 0.408 | 0.441 | 1,981 |
The 9070 XT's render time (0.408 ms) is 3× that of the 6900 XT (0.134 ms), despite rendering the same single draw call of point sprites. Investigation revealed this is not a GPU performance issue but a measurement artefact caused by swapchain semaphore wait pollution in timestamps.
In the Vulkan backend, timestamp T3 is written at VK_PIPELINE_STAGE_COLOR_ATTACHMENT_OUTPUT_BIT. This pipeline stage does not begin until the presentation engine releases a swapchain image — signalled via the imageAvailableSemaphore acquired from vkAcquireNextImageKHR.
T2 (TOP_OF_PIPE) ──→ Vertex Shader ──→ Rasterisation ──→ Fragment Shader
│
┌─────────────┘
▼
COLOR_ATTACHMENT_OUTPUT
┌──────────────────────┐
│ Wait for semaphore │ ← swapchain image availability
│ (variable delay) │
│ Actual pixel write │
└──────────────────────┘
│
▼
T3 (timestamp)
The render time (T3 − T2) therefore measures: actual render work + semaphore wait for swapchain image. On fast GPUs that finish compute + render in under 1 ms, the GPU spends most of its time idle, waiting for the presentation engine to recycle a swapchain image from the previous frame's vkQueuePresentKHR.
| GPU | Situation | Effect on Render Timestamp |
|---|---|---|
| RTX 5090 | CPU-bound (30% GPU util). GPU finishes early, waits for next CPU submission. Semaphore wait is hidden within CPU stall. | Minimal pollution — GPU idle time is between frames, not during render stage |
| RX 6900 XT | Balanced. Compute takes long enough (0.063 ms) that the presentation engine has time to release the swapchain image before render stage begins. | Low pollution — semaphore is usually already signalled |
| RX 9070 XT | Compute finishes very fast (0.033 ms), reaches COLOR_ATTACHMENT_OUTPUT before the presentation engine has released the image from the previous present. GPU stalls waiting for semaphore. |
High pollution — ~0.37 ms of the 0.408 ms "render time" is semaphore wait |
The 9070 XT's actual render work is approximately 0.04 ms (similar to other GPUs), but the timestamp reports 0.408 ms because it includes the swapchain image wait.
Note that at 1M particles, the render timestamp difference is primarily a semaphore wait artefact, not a real GPU performance gap. However, at higher particle counts (16M), the render time difference becomes real — and Infinity Cache becomes a significant factor: the 6900 XT's 128 MB cache vs the 9070 XT's 64 MB provides a substantial hit-rate advantage for primitive-heavy rendering (see Section 5b for full analysis).
Two frequently confused concepts that both affect frame pacing:
| Swapchain BufferCount | VSync | |
|---|---|---|
| What it controls | How many swapchain images exist in the pool | When completed frames are shown on the display |
| Set by | VkSwapchainCreateInfoKHR::minImageCount (Vulkan), DXGI_SWAP_CHAIN_DESC::BufferCount (DX11/12) |
glfwSwapInterval() (OpenGL), Present(syncInterval) (DX), VK_PRESENT_MODE_* (Vulkan) |
| Effect on FPS | More buffers → less GPU idle time waiting for image availability, but diminishing returns beyond 3 | VSync ON → FPS capped to display refresh rate; VSync OFF → uncapped |
| Effect on input lag | More buffers → higher input lag (more pre-rendered frames queued) | VSync ON → adds up to one frame of latency |
Increasing BufferCount from 2 to 3 was tested (via --flights 3) but showed no meaningful improvement in timestamp pollution. The semaphore wait is inherent to the Vulkan presentation model — the timestamp at COLOR_ATTACHMENT_OUTPUT_BIT will always include the wait, regardless of how many images are in the pool, because the GPU must still wait for at least one image to become available from the presentation engine.
To measure pure GPU compute performance without any swapchain, rendering, or presentation interference, a headless compute mode (--headless) was implemented:
| Component | Windowed Mode | Headless Mode |
|---|---|---|
| Window | GLFW visible window | No window (OpenGL: hidden window for context) |
| Swapchain | Created, images acquired/presented | Not created |
| Render pass | Full vertex + fragment pipeline | Skipped entirely |
| Present | vkQueuePresentKHR / Present() / glfwSwapBuffers |
Skipped |
| Timestamps | T0–T3 (compute + render) | T0–T1 (compute only), T2=T3 mirrored |
| GPU utilisation | Limited by presentation engine | Limited only by compute throughput |
| API | FPS | Compute (ms) | GPU Util |
|---|---|---|---|
| Vulkan | 21,260 | 0.034 | 70% |
| DX12 | 21,354 | 0.034 | 70% |
| DX11 | 16,938 | 0.034 | 60% |
| OpenGL | 20,298 | 0.034 | 70% |
All four APIs converge to nearly identical compute times (0.034 ms), confirming that the GPU-side compute workload is equivalent across APIs. The FPS difference reflects only CPU-side overhead — DX11 is slightly slower due to its implicit driver model requiring more CPU work per frame without a Present() call to batch around.
Compare with windowed mode:
| API | Windowed FPS | Headless FPS | Speedup |
|---|---|---|---|
| Vulkan | 1,981 | 21,260 | 10.7× |
| DX12 | 2,100 | 21,354 | 10.2× |
| DX11 | 3,500 | 16,938 | 4.8× |
| OpenGL | 1,800 | 20,298 | 11.3× |
The 10× speedup confirms that windowed mode performance is dominated by presentation overhead, not by compute or render workload.
DX11's implicit driver model uses Present() as an implicit frame boundary for command batching and timestamp query resolution. Without it:
- Problem 1:
CollectTimestampResults()usedSleep(1)retries waiting for query resolution. Without Present(), queries never resolved promptly, causing each frame to take 4+ ms (Sleep granularity). - Fix: Removed Sleep in headless mode; spin-wait only.
- Problem 2: Even with spin-wait, timestamp values were occasionally garbage (e.g., 805534675707 ms) because DX11 lacks proper frame boundaries without Present().
- Fix: Added
context_->Flush()after compute dispatch to force command submission, plus sanity filter discarding timestamps > 1000 ms. Approximately 3–4% of frames produce garbage values and are discarded (e.g., 212531/220193 valid samples).
OpenGL: AMD Driver Requires Explicit Flush for Hidden Windows
OpenGL requires a window (even hidden) to create a GL context. On AMD drivers, glFlush() alone is insufficient to process commands for hidden windows — the driver does not actively schedule GPU work without a visible surface.
- Attempt 1:
glFlush()only → timestamps never resolve (0 valid samples). - Attempt 2:
glFinish()every frame → timestamps work, but FPS drops from 22,000 to 8,000 (CPU stalls waiting for GPU). - Final solution:
glFinish()every 16th frame (forces command processing) +glFenceSync+glFlush()on other frames (non-blocking). This achieves 20,298 FPS with ~25% timestamp sample rate (65985/263889 valid samples).
Both explicit APIs handle headless cleanly:
- Skip swapchain, render pass, and present calls
- Compute dispatch + fence sync is sufficient
- 100% timestamp sample rate, no workarounds needed
3DMark offers an Unlimited mode that removes VSync and frame rate caps. This is often confused with headless compute, but they are fundamentally different:
| 3DMark Unlimited | This Benchmark Headless | |
|---|---|---|
| Rendering | Full offscreen rendering (all geometry, textures, post-FX) | No rendering — compute dispatch only |
| Target | Offscreen render target (no swapchain present) | No render target at all |
| Measures | Combined compute + render + post-processing GPU throughput, uncapped | Pure compute shader throughput |
| Presentation | Skipped (no VSync, no Present) | Skipped |
| Use case | Cross-device comparison without display refresh rate bias | Isolating compute performance from presentation overhead |
| Analogy | Running the full game engine but rendering to a texture instead of screen | Running only the physics engine with no rendering at all |
3DMark Unlimited is equivalent to rendering to an offscreen framebuffer — the full GPU pipeline (vertex → rasterisation → fragment → post-processing) executes, but the final present/flip is skipped. This benchmark's headless mode is more aggressive: it eliminates the entire graphics pipeline, measuring only the compute dispatch that updates particle positions.
If a 3DMark Unlimited-style mode were added to this benchmark, it would involve:
- Creating an offscreen framebuffer (VkFramebuffer / ID3D11RenderTargetView / FBO)
- Running the full compute + render pipeline to that framebuffer
- Skipping only
vkQueuePresentKHR/Present()/glfwSwapBuffers - Timestamp T3 would measure actual render completion without semaphore wait pollution
This would provide a middle ground between windowed (presentation-throttled) and headless (compute-only) modes, and would be the most direct comparison point with 3DMark Unlimited scores.
Cross-generational compute shader performance comparison across eight AMD GPUs spanning five architectures and 16 years of hardware evolution (2009–2025). All results collected with this project's particle simulation benchmark.
| GPU | Architecture | Generation | CUs / SPs | Core Clock | FP32 TFLOPS | Memory | Bandwidth | Platform | API Coverage |
|---|---|---|---|---|---|---|---|---|---|
| HD 5770 | TeraScale 2 | 2009 | 800 SPs (VLIW5) | 850 MHz | ~1.36 | 1 GB GDDR5 | 76.8 GB/s | Windows | DX11, OpenGL |
| FirePro D700 | GCN 1.0 (Tahiti) | 2013 | 2048 SPs | 850 MHz | ~3.5 | 6 GB GDDR5 | 264 GB/s | Windows | Vulkan, DX12, DX11, OpenGL |
| RX 580 | GCN 4 (Polaris) | 2017 | 36 CUs | 1,340 MHz | ~6.2 | 8 GB GDDR5 | 256 GB/s | Windows | Vulkan, DX12, DX11, OpenGL |
| Vega Frontier Edition | GCN 5 (Vega) | 2017 | 64 CUs | 1,600 MHz | ~13.1 | 16 GB HBM2 | 483 GB/s | Windows | Vulkan, DX12, DX11, OpenGL |
| RX 6600 XT | RDNA 2 | 2021 | 32 CUs | 2,589 MHz | ~10.6 | 8 GB GDDR6 | 256 GB/s | Windows | Vulkan, DX12, DX11, OpenGL |
| RX 6900 XT | RDNA 2 | 2020 | 80 CUs | 2,250 MHz | ~23.0 | 16 GB GDDR6 | 512 GB/s | Windows | Vulkan, DX12, DX11, OpenGL |
| RX 9070 XT | RDNA 4 | 2025 | 64 CUs | 2,970 MHz | ~48.7 † | 16 GB GDDR6 | 640 GB/s | Windows | Vulkan, DX12, DX11, OpenGL |
| Ryzen 9800X3D iGPU | RDNA 2 | 2024 | 2 CUs | 2,200 MHz | ~0.56 | Shared DDR5 | ~83 GB/s | Windows | Vulkan, DX12, DX11, OpenGL |
Note: FirePro D700 data is from Windows (Boot Camp). Vulkan, DX12, DX11, and OpenGL results are available for GCN 1.0. All GPUs are tested on Windows. † RDNA 4 uses dual-issue FP32 (each SP executes 2 FP32 ops/clock); traditional single-issue calculation yields ~24.3 TFLOPS.
Best API per GPU (highest FPS):
| # | GPU | Architecture | Best API | Avg FPS | Compute (ms) | Render (ms) | Total GPU (ms) | Bottleneck |
|---|---|---|---|---|---|---|---|---|
| 1 | RX 9070 XT | RDNA 4 (64 CU) | DX11 | 1,773.7 | 0.047 | 0.451 | 0.542 | GPU-bound |
| 2 | RX 6900 XT | RDNA 2 (80 CU) | DX11 | 4,067.9 | 0.047 | 0.140 | 0.226 | GPU-bound |
| 3 | RX 6600 XT | RDNA 2 (32 CU) | DX12 | 1,834.4 | 0.190 | 0.215 | 0.406 | Balanced |
| 3 | Vega FE | GCN 5 (64 CU) | DX12 | 1,715.5 | 0.219 | 0.229 | 0.452 | Balanced |
| 4 | RX 580 | GCN 4 (36 CU) | DX12 | 912.2 | 0.362 | 0.557 | 0.930 | GPU-bound |
| 5 | FirePro D700 | GCN 1.0 | Vulkan | 554.8 | 0.589 | 0.881 | 1.473 | GPU-bound |
| 6 | Zen4/5 iGPU (2 CU) | RDNA 2 | DX12 | 324.0 | 1.480 | 1.472 | 2.953 | GPU-bound |
| 7 | HD 5770 | TeraScale 2 | OpenGL | 188.2 | 1.794 | 3.018 | 4.818 | GPU-bound |
| 8 | WARP on 9800X3D¹ | Software | DX12 | 86.6 | 1.034 | 10.024 | 11.059 | Software |
| 9 | WARP on 7600¹ | Software | DX12 | 62.2 | 1.810 | 13.369 | 15.181 | Software |
¹ WARP runs on the CPU, not a GPU. The two WARP entries show the same software renderer on different CPUs: the Ryzen 7 9800X3D (8-core, 96 MB L3 3D V-Cache) is 39% faster than the Ryzen 5 7600 (6-core, 32 MB L3), demonstrating how CPU core count, clock speed, and cache size directly affect software rendering performance.
All API results per GPU:
| GPU | Vulkan | DX12 | DX11 | OpenGL | Metal |
|---|---|---|---|---|---|
| RX 9070 XT | 1,751 FPS | 1,609 FPS | 1,774 FPS | 253 FPS | N/A |
| RX 6900 XT | 2,885 FPS | 3,518 FPS | 4,068 FPS | 229 FPS | N/A |
| RX 6600 XT | 1,239 FPS | 1,834 FPS | 988 FPS | 180 FPS | N/A |
| Vega FE | 1,370 FPS | 1,716 FPS | 1,436 FPS | 158 FPS | N/A |
| RX 580 | 783 FPS | 912 FPS | 755 FPS | 42 FPS | N/A |
| FirePro D700 | 555 FPS | 516 FPS | 527 FPS | 525 FPS | N/A |
| Zen4/5 iGPU (2 CU) | 238 FPS | 324 FPS | 271 FPS | 229 FPS | N/A |
| HD 5770 | N/A | N/A | 107 FPS | 188 FPS | N/A |
Normalise compute shader performance to per-CU throughput to isolate architectural efficiency from raw CU count.
| GPU | Architecture | CUs | Compute Time (ms) | Per-CU Throughput (relative) | Per-CU vs RX 580 |
|---|---|---|---|---|---|
| HD 5770 | TeraScale 2 | ~10 equiv | 1.794 (OpenGL) | 0.0557 | 0.73× |
| FirePro D700 | GCN 1.0 | 32 | 0.589 | 0.0531 | 0.69× |
| RX 580 | GCN 4 | 36 | 0.362 | 0.0767 | 1.00× |
| Vega FE | GCN 5 | 64 | 0.219 | 0.0713 | 0.93× |
| RX 6600 XT | RDNA 2 | 32 | 0.270 | 0.1157 | 1.51× |
| RX 9070 XT | RDNA 4 | 64 | 0.033 | 0.4735 | 6.17× |
| RX 6900 XT | RDNA 2 | 80 | 0.063 | 0.1984 | 2.59× |
| Zen4/5 iGPU (2 CU) | RDNA 2 | 2 | 1.257 | 0.3978 | 5.19× |
HD 5770 CU equivalence: TeraScale 2 does not have CUs. 800 VLIW5 stream processors are roughly grouped into 10 SIMD engines. This mapping is approximate.
Analysis:
Per-CU efficiency improves modestly from GCN 1.0 (0.69x) through GCN 4 (1.00x) to GCN 5 (0.93x), with Vega FE slightly below RX 580 per-CU despite being a newer architecture -- likely because the 64 CUs are not fully utilised at 1M particles. The jump to RDNA 2 brings a 1.51x per-CU improvement on the RX 6600 XT, confirming the architectural efficiency gain from GCN to RDNA.
The standout result is RDNA 4: the RX 9070 XT achieves 6.17× the per-CU throughput of the RX 580, a significant generational leap. This is partly real architectural improvement (higher clocks, better scheduler, 640 GB/s bandwidth vs 256 GB/s) and partly because the 0.033 ms compute time is near the floor of per-dispatch overhead, meaning the GPU finishes so quickly that fixed overhead dominates less.
The iGPU (2 CU) shows 5.19x per-CU efficiency vs the RX 580, which is surprisingly high. With only 2 CUs, the workgroup scheduling overhead is minimal and the entire working set fits in cache, inflating per-CU throughput. The RX 6900 XT (2.59x) has lower per-CU than the iGPU despite being the same RDNA 2 architecture, confirming that CU scaling is sub-linear -- more CUs means more contention for memory bandwidth and cache.
Three RDNA 2 GPUs at vastly different CU counts allow direct measurement of how compute performance scales with CU count within the same architecture.
| GPU | CUs | FPS | Compute (ms) | Scaling vs iGPU (2 CU) | Ideal Scaling (CU ratio) | Efficiency |
|---|---|---|---|---|---|---|
| Zen4/5 iGPU (2 CU) | 2 | 238 | 1.257 | 1.00× | 1.00× | 100% |
| RX 6600 XT | 32 | 1,239 | 0.270 | 4.66× | 16.0× | 29% |
| RX 6900 XT | 80 | 2,885 | 0.063 | 19.95× | 40.0× | 50% |
Analysis:
RDNA 2 CU scaling is heavily sub-linear at 1M particles. The RX 6600 XT (16x the CUs) achieves only 4.66x the compute throughput of the iGPU, a 29% scaling efficiency. The RX 6900 XT (40x the CUs) reaches 19.95x, or 50% efficiency -- better than the 6600 XT because its 512 GB/s bandwidth (vs 256 GB/s) helps feed the additional CUs.
The primary bottleneck is memory bandwidth. At 1M particles (32 MB SSBO), the working set is small enough to benefit from cache on the iGPU but must stream through GDDR6 on the discrete cards. The iGPU's DDR5 bandwidth (~83 GB/s) is fully utilised by just 2 CUs, while the RX 6600 XT's 256 GB/s must be shared across 32 CUs. At higher particle counts (16M+), scaling efficiency would improve as the per-CU workload increases and fixed dispatch overhead is amortised.
Normalise all GPUs to the RX 580 (GCN 4) = 1.00× baseline for generational comparison.
Windowed 1M particles (presentation-limited for fast GPUs):
| GPU | Architecture | Year | FPS | vs RX 580 | Memory BW | BW vs RX 580 |
|---|---|---|---|---|---|---|
| HD 5770 | TeraScale 2 | 2009 | 188 | 0.21× | 76.8 GB/s | 0.30× |
| FirePro D700 | GCN 1.0 | 2013 | 555 | 0.61× | 264 GB/s | 1.03× |
| RX 580 | GCN 4 | 2017 | 912 | 1.00× | 256 GB/s | 1.00× |
| Vega FE | GCN 5 | 2017 | 1,716 | 1.88× | 483 GB/s | 1.89× |
| RX 6600 XT | RDNA 2 | 2021 | 1,834 | 2.01× | 256 GB/s | 1.00× |
| RX 6900 XT | RDNA 2 | 2020 | 4,068 | 4.46× | 512 GB/s | 2.00× |
| RX 9070 XT | RDNA 4 | 2025 | 1,774 | 1.95× | 640 GB/s | 2.50× |
Headless 1M particles (pure compute, no presentation overhead):
| GPU | Architecture | Headless FPS (best API) | vs RX 580 | Windowed → Headless Speedup |
|---|---|---|---|---|
| RX 6900 XT | RDNA 2 (80 CU) | 15,950 | 17.49× | 3.9× |
| RX 9070 XT | RDNA 4 (64 CU) | 21,354 | 23.41× | 12.0× |
Analysis:
The windowed 1M table has an inherent limitation: fast GPUs are presentation-limited, so FPS reflects swapchain throughput rather than GPU compute speed. The RX 9070 XT (1.95×) appears slower than the RX 6900 XT (4.46×) despite having faster per-CU compute — because swapchain throttling caps the 9070 XT more severely (its 64 CUs finish compute faster, spending more time waiting for presentation).
The headless data removes this distortion. With presentation overhead eliminated:
- The RX 9070 XT is the fastest AMD GPU tested at 21,354 FPS — 1.34× faster than the 80-CU RX 6900 XT (15,950 FPS), despite having only 80% of its CU count. This means RDNA 4's per-CU compute efficiency is ~1.7× higher than RDNA 2.
- The windowed → headless speedup reveals how much performance is hidden by presentation: the 9070 XT unlocks 12× more throughput in headless mode, while the 6900 XT unlocks only 3.9× — because the 6900 XT's 80 CUs already partially saturate the presentation pipeline in windowed mode.
Bandwidth remains the dominant performance predictor on GCN: Vega FE achieves 1.88× the FPS of the RX 580, almost exactly matching its 1.89× bandwidth ratio. The RX 6600 XT breaks this pattern: identical bandwidth to the RX 580 (256 GB/s, 1.00×) yet 2.01× the FPS — evidence that RDNA 2's architectural improvements (better cache, improved scheduler, higher per-CU throughput) contribute beyond raw bandwidth.
TFLOPS remains a poor predictor: Vega FE has 13.1 TFLOPS (2.1× the RX 580's 6.2) but only 1.88× FPS, while the RX 6600 XT has 10.6 TFLOPS (1.7×) but 2.01× FPS. This workload is bandwidth-bound, not compute-bound.
Does the API ranking (DX11 > DX12 > Vulkan > OpenGL) observed on RTX 5090 hold across all AMD architectures, or does it change?
| GPU | Architecture | Fastest API | DX11 FPS | DX12 FPS | Vulkan FPS | OpenGL FPS | DX11 vs Vulkan Gap |
|---|---|---|---|---|---|---|---|
| HD 5770 | TeraScale 2 | OpenGL | 107 | N/A | N/A | 188 | N/A |
| RX 580 | GCN 4 | DX12 | 755 | 912 | 783 | 42 | −3.6% |
| Vega FE | GCN 5 | DX12 | 1,436 | 1,716 | 1,370 | 158 | +4.8% |
| RX 6600 XT | RDNA 2 | DX12 | 988 | 1,834 | 1,239 | 180 | −20.3% |
| RX 9070 XT | RDNA 4 | DX11 | 1,774 | 1,609 | 1,751 | 253 | +1.3% |
| RX 6900 XT | RDNA 2 | DX11 | 4,068 | 3,518 | 2,885 | 229 | +41.0% |
| Zen4/5 iGPU (2 CU) | RDNA 2 | DX12 | 271 | 324 | 238 | 229 | +13.9% |
Analysis:
Unlike the RTX 5090 where DX11 was consistently fastest, AMD GPUs show a more varied API ranking. DX12 is the fastest API on four of seven GPUs (RX 580, Vega FE, RX 6600 XT, iGPU), while DX11 wins on the RX 9070 XT and RX 6900 XT. This suggests AMD's DX12 driver is more competitive than NVIDIA's for simple workloads, while DX11 still benefits the fastest GPUs where CPU overhead matters most.
The DX11 vs Vulkan gap is much smaller on AMD than on NVIDIA. The RX 9070 XT shows only +1.3% for DX11 over Vulkan, and the RX 580 shows just -3.6% (Vulkan slightly faster). The largest gap is the RX 6900 XT at +41.0%, where DX11 dramatically outperforms Vulkan -- likely because at 4,068 FPS, every microsecond of per-frame CPU overhead matters, and AMD's DX11 driver path has lower latency than the Vulkan submission path at this extreme frame rate.
The RX 6600 XT shows a notable -20.3% gap (Vulkan faster than DX11), indicating that AMD's DX11 driver for RDNA 2 mid-range parts has higher overhead than the Vulkan path. This contrasts with the 6900 XT result and suggests driver optimisation varies by SKU.
OpenGL is not competitive on any modern AMD GPU, falling to 42 FPS on the RX 580 (21x slower than DX12). The sole exception is the HD 5770 where OpenGL (188 FPS) outperforms DX11 (107 FPS) -- TeraScale 2 predates compute shaders in DX11, and the OpenGL path may use a more efficient fallback. The FirePro D700 is notable for near-identical performance across all four APIs (516-555 FPS), suggesting its GCN 1.0 architecture is purely GPU-bound regardless of API overhead.
Plot FPS against memory bandwidth to test the hypothesis that this benchmark is bandwidth-bound.
| GPU | Memory BW (GB/s) | FPS (best API) | FPS / (GB/s) |
|---|---|---|---|
| HD 5770 | 76.8 | 188 | 2.45 |
| FirePro D700 | 264 | 555 | 2.10 |
| RX 580 | 256 | 912 | 3.56 |
| Vega FE (HBM2) | 483 | 1,716 | 3.55 |
| RX 6600 XT | 256 | 1,834 | 7.16 |
| RX 6900 XT | 512 | 4,068 | 7.95 |
| RX 9070 XT | 640 | 1,774 | 2.77 |
| Zen4/5 iGPU (DDR5) | ~83 | 324 | 3.90 |
Analysis:
The FPS/(GB/s) ratio reveals two distinct tiers of bandwidth efficiency. Older architectures (HD 5770, FirePro D700, RX 580, Vega FE) cluster around 2.1-3.6 FPS per GB/s, while RDNA 2 parts (RX 6600 XT, RX 6900 XT) achieve 7.2-8.0 FPS per GB/s -- roughly 2x the bandwidth efficiency. This confirms that RDNA 2's Infinity Cache dramatically amplifies effective bandwidth for workloads with temporal locality, as the 32 MB SSBO working set partially fits in the cache.
Vega FE (3.55) matches RX 580 (3.56) almost exactly in FPS per GB/s despite being a different architecture (GCN 5 vs GCN 4). This means Vega FE's 1.88x FPS advantage comes almost entirely from its 1.89x bandwidth advantage (HBM2), with negligible architectural efficiency gain for this workload.
The RX 9070 XT (3.46) falls to the GCN-era efficiency tier despite being RDNA 4, because its FPS is presentation-limited at 1,774 FPS. The GPU finishes compute in 0.033 ms but waits for swapchain presentation, wasting most of its bandwidth potential.
The iGPU (3.90) performs slightly above the GCN tier despite using shared DDR5. Its 2 CUs generate low enough memory traffic that DDR5 bandwidth is not contended with CPU traffic, and the unified memory architecture avoids PCIe overhead entirely.
Run each GPU at multiple particle counts to find the crossover point where the bottleneck shifts from CPU to GPU.
| GPU | Architecture | 1M FPS (best API) | 16M FPS (best API) | 16×/1M Ratio |
|---|---|---|---|---|
| RTX 5090 | Blackwell (170 SM) | 7,737 | 612 | 12.6× |
| RX 9070 XT | RDNA 4 (64 CU) | 1,774 | 111 | 16.0× |
| RX 6900 XT | RDNA 2 (80 CU) | 4,068 | 218 | 18.7× |
| RX 6600 XT | RDNA 2 (32 CU) | 1,834 | — | — |
| Vega FE | GCN 5 (64 CU) | 1,716 | — | — |
| RX 580 | GCN 4 (36 CU) | 912 | — | — |
| HD 5770 | TeraScale 2 | 188 | — | — |
| Zen4/5 iGPU (2 CU) | RDNA 2 | 324 | — | — |
Note: The RTX 5090 ($1,999, flagship tier) is included as a cross-vendor reference point, not as a direct competitor to the mid-range RX 9070 XT ($599) or RX 6900 XT (launched $999, now ~$400 used). Price-performance analysis is in Section 8.
Analysis:
At 1M particles, the fast GPUs are presentation-limited — FPS reflects swapchain throughput rather than GPU compute speed. The 16M column removes this limitation and reveals true GPU-bound performance.
- RX 6900 XT leads among AMD GPUs at 16M (218 FPS vs 111 FPS for the 9070 XT). This is not a compute advantage — the 9070 XT is faster in pure compute (headless: 21K vs 15K FPS). The difference is the render pass: 16M point primitives stress raw rasterisation throughput, where the 6900 XT's 80 CUs provide 1.25× more parallel rasterisation hardware than the 9070 XT's 64 CUs (detailed analysis in Section 5b).
- Scaling ratios vary by architecture: the RX 6900 XT's 18.7× ratio (vs expected 16×) reflects its 1M score being more presentation-limited than the 9070 XT's. The RTX 5090's 12.6× suggests its 1M score was already partially GPU-bound (less room to "unlock").
- RTX 5090 at 16M (612 FPS) is 2.8× the RX 6900 XT — a substantial gap, but narrower than the ~4× implied by their TFLOPS ratio (105 vs 23 TFLOPS), confirming this workload is bandwidth-bound rather than compute-bound.
Capture one Vulkan frame on each AMD GPU (where Vulkan is available) and compare per-event GPU timing, barrier cost, and command structure.
Captures generated via
--capture 5(auto-capture at 5 seconds). Analysis automated withscripts/rdoc_analyse.pyandscripts/rdoc_export_timing.py.
| GPU | Architecture | Compute Dispatch (ms) | Barrier (ms) | Render Pass (ms) | Total GPU (ms) |
|---|---|---|---|---|---|
| RX 9070 XT | RDNA 4 (64 CU) | 0.033 | 0.005 | 0.408 | 0.446 |
| RX 6900 XT | RDNA 2 (80 CU) | 0.063 | < 0.001 | 0.133 | 0.196 |
| RX 6600 XT | RDNA 2 (32 CU) | 0.270 | < 0.001 | 0.379 | 0.649 |
| Vega FE | GCN 5 (64 CU) | 0.219 | 0.001 | 0.275 | 0.495 |
| RX 580 | GCN 4 (36 CU) | 0.362 | 0.006 | 0.702 | 1.070 |
| FirePro D700 | GCN 1.0 (12 CU) | 0.589 | 0.003 | 0.881 | 1.473 |
| Zen4/5 iGPU (2 CU) | RDNA 2 | 1.257 | < 0.001 | 1.928 | 3.185 |
HD 5770 excluded — no Vulkan support. Barrier cost derived from
Total GPU − Compute − Render(timestamps 1→4 minus timestamps 1→2 and 3→4). Values < 0.001 ms are below timestamp query resolution.
| GPU | Metric | App Timestamp (ms) | RenderDoc CPU Trace (µs) | Notes |
|---|---|---|---|---|
| RX 9070 XT | Compute | 0.033 | 0.009 | CPU recording ≪ GPU execution |
| RX 9070 XT | Render | 0.408 | 0.051 | GPU-bound render pass |
| RX 6900 XT | Compute | 0.063 | 0.004 | CPU recording ≪ GPU execution |
| RX 6900 XT | Render | 0.133 | 0.074 | GPU-bound render pass |
| RX 6600 XT | Compute | 0.270 | 0.004 | CPU recording ≪ GPU execution |
| RX 6600 XT | Render | 0.379 | 0.076 | GPU-bound render pass |
| Vega FE | Compute | 0.219 | 0.005 | CPU recording ≪ GPU execution |
| Vega FE | Render | 0.275 | 0.099 | GPU-bound render pass |
| RX 580 | Compute | 0.362 | 0.004 | CPU recording ≪ GPU execution |
| RX 580 | Render | 0.702 | 0.075 | GPU-bound render pass |
| Zen4/5 iGPU (2 CU) | Compute | 1.257 | 0.005 | CPU recording ≪ GPU execution |
| Zen4/5 iGPU (2 CU) | Render | 1.928 | 0.107 | GPU-bound render pass |
Note: RenderDoc JSON exports record CPU-side API call durations (when each
vkCmd*was recorded into the command buffer), not GPU execution time. GPU-side per-event timing requires opening the.rdcfiles in the RenderDoc GUI Performance Counter Viewer. The CPU trace confirms all GPUs are fully GPU-bound: CPU recording is 100–1000× faster than GPU execution for every event.
| GPU | Architecture | Memory Type | Barrier Duration (ms) | Notes |
|---|---|---|---|---|
| RX 9070 XT | RDNA 4 | 16 GB GDDR6 | 0.005 | Measurable — RDNA 4 L2 writeback |
| RX 6900 XT | RDNA 2 | 16 GB GDDR6 | < 0.001 | Below timestamp resolution |
| RX 6600 XT | RDNA 2 | 8 GB GDDR6 | < 0.001 | Below timestamp resolution |
| Vega FE | GCN 5 | 16 GB HBM2 | 0.001 | HBM2 — near-zero |
| RX 580 | GCN 4 | 8 GB GDDR5 | 0.006 | Highest — GDDR5 L2 flush |
| FirePro D700 | GCN 1.0 | 6 GB GDDR5 | 0.003 | Moderate — early GCN |
| Zen4/5 iGPU (2 CU) | RDNA 2 | Shared DDR5 | < 0.001 | Unified memory — near-zero |
Analysis:
Barrier cost is negligible across all tested AMD GPUs (< 0.006 ms), confirming that the VK_PIPELINE_STAGE_COMPUTE_SHADER_BIT → VK_PIPELINE_STAGE_VERTEX_INPUT_BIT buffer memory barrier introduces minimal synchronisation overhead. The RX 580 (GCN 4, GDDR5) shows the highest measurable barrier at 0.006 ms, likely due to its older L2 cache architecture requiring a full writeback. RDNA 2 GPUs (6900 XT, 6600 XT, iGPU) all show sub-microsecond barriers, suggesting the Infinity Cache absorbs the coherency cost. The iGPU's unified memory architecture confirms the expected near-zero barrier cost — no physical cache flush is needed when compute and graphics share the same memory controller.
| GPU | Total Events | Frame Events | Dispatches | Draw Calls | Barriers | Driver-Inserted |
|---|---|---|---|---|---|---|
| RX 9070 XT | ~135 | 52 | 1 | 1 | 1 | 0 |
| RX 6900 XT | ~130 | 50 | 1 | 1 | 1 | 0 |
| RX 6600 XT | ~135 | 51 | 1 | 1 | 1 | 0 |
| Vega FE | ~130 | 52 | 1 | 1 | 1 | 0 |
| RX 580 | ~130 | 50 | 1 | 1 | 1 | 0 |
| FirePro D700 | ~130 | 50 | 1 | 1 | 1 | 0 |
| Zen4/5 iGPU (2 CU) | ~135 | 52 | 1 | 1 | 1 | 0 |
All AMD GPUs produce an identical frame structure: 1 compute dispatch, 1 pipeline barrier, 1 render pass (begin + draw + end), 4 timestamp writes, and 3 debug label pairs. Frame event counts (50–52) vary slightly due to some debug label events being recorded as instant events vs begin/end pairs across driver versions. No driver-inserted implicit barriers observed on any AMD GPU — the single explicit
vkCmdPipelineBarrieris sufficient.NVIDIA RTX 5090 comparison (Section 3): 133 total events, 28 frame events — the higher total reflects NVIDIA's driver inserting additional internal resource tracking events outside the frame boundary.
Cross-validate this benchmark's AMD GPU rankings against 3DMark Time Spy (DX12) and Fire Strike (DX11) to confirm the results reflect real-world performance scaling.
Data source:
scripts/3dmark_scores.jsonCharts generated with:python scripts/compare_3dmark.py --save docs/images
| GPU | Architecture | This Benchmark | 3DMark Time Spy | 3DMark Fire Strike | Deviation (TS) | Deviation (FS) |
|---|---|---|---|---|---|---|
| RX 9070 XT | RDNA 4 (64 CU) | 1.95× | 6.49× | 4.44× | −70.0% | −56.1% |
| RX 6900 XT | RDNA 2 (80 CU) | 4.46× | 4.63× | 3.82× | −3.7% | +16.8% |
| RX 6600 XT | RDNA 2 (32 CU) | 2.01× | 2.16× | 1.92× | −7.0% | +4.9% |
| Vega FE | GCN 5 (64 CU) | 1.88× | 1.59× | 1.46× | +18.2% | +28.8% |
| RX 580 | GCN 4 (36 CU) | 1.00× | 1.00× | 1.00× | — | — |
| Zen4/5 iGPU (2 CU) | RDNA 2 | 0.36× | 0.16× | 0.15× | +119.2% | +132.0% |
| HD 5770 | TeraScale 2 | 0.21× | N/A | 0.10× | N/A | +106.8% |
Deviation =
(This Benchmark ratio / 3DMark ratio) − 1. Positive = our benchmark favours that GPU more; negative = less.
| GPU | Expected Deviation | Reason |
|---|---|---|
| RX 9070 XT | Negative (−50–70%) | FPS is presentation-limited at 1,774 FPS; 3DMark exercises full GPU feature set |
| RX 6900 XT | Near zero (TS) / Positive (FS) | 80 CU + 512 GB/s bandwidth scales well for both workloads; Fire Strike's DX11 overhead less efficient than Vulkan compute |
| RX 6600 XT | Near zero | Mid-range GPU; balanced for both workload types |
| Vega FE | Positive (+18–29%) | HBM2's 483 GB/s bandwidth disproportionately benefits bandwidth-bound compute workloads |
| Zen4/5 iGPU (2 CU) | Positive (+119–132%) | Simple compute fits within 2-CU cache; 3DMark's complex workloads expose shader count limit |
| HD 5770 | N/A for Time Spy | TeraScale 2 has no DX12; Fire Strike deviation +107% (similar to iGPU — simple compute overperforms) |
The deviations reveal how this benchmark's single-dispatch compute workload differs from 3DMark's complex multi-pass rasterisation:
- RX 9070 XT (−70.0% TS, −56.1% FS): The largest deviation. In 3DMark, the 9070 XT exercises its full RDNA 4 feature set (mesh shaders, ray tracing units, 64 CUs at full utilisation). In this benchmark, windowed FPS is presentation-limited at 1,774 FPS — the GPU finishes compute in 0.033 ms but waits for swapchain presentation (Section 6). Headless mode resolves this: at 21,354 FPS (23.41× vs RX 580), the deviation flips to +261% (TS) and +427% (FS) — confirming that when presentation overhead is removed, this compute benchmark massively favours the 9070 XT over 3DMark's rasterisation workloads. The windowed deviation measures presentation bottleneck; the headless deviation measures the fundamental compute-vs-rasterisation gap.
- Vega FE (+18.2% TS, +28.8% FS): HBM2's 483 GB/s bandwidth disproportionately benefits this bandwidth-bound compute workload. 3DMark's texture-heavy scenes do not benefit as much from raw bandwidth.
- iGPU (+119% TS, +132% FS): The 2-CU iGPU overperforms relative to 3DMark because this benchmark's simple compute workload fits well within the iGPU's cache and DDR5 bandwidth is adequate for 2 CUs. 3DMark's complex geometry and texture workloads expose the iGPU's limited shader count.
- RX 6900 XT (−3.7% TS) and RX 6600 XT (−7.0% TS): Within the expected ±10% range for Time Spy in windowed mode, confirming that RDNA 2 GPUs are well-correlated. In headless mode, the 6900 XT jumps to 17.49× (vs 4.46× windowed), yielding +278% (TS) and +358% (FS) deviations — similar to the 9070 XT pattern, confirming that all fast GPUs are presentation-limited in windowed mode. Fire Strike deviations are slightly higher (+16.8% and +4.9%) because DX11's heavier CPU overhead in 3DMark penalises complex scenes more than this benchmark's single-dispatch Vulkan path.
- HD 5770 (+106.8% FS): Similar to the iGPU — TeraScale 2's limited shader hardware handles this simple compute workload relatively better than 3DMark's complex rasterisation. No Time Spy comparison possible (no DX12).
The large deviations for the 9070 XT (presentation-limited) and the low-end GPUs (iGPU, HD 5770) confirm that this benchmark measures a fundamentally different aspect of GPU performance (bandwidth-bound compute + presentation overhead) compared to 3DMark (shader-heavy rasterisation). GPUs in their GPU-bound regime (6900 XT, 6600 XT) show strong correlation. Both benchmarks are needed for a complete performance picture.
| This Benchmark | 3DMark Time Spy | 3DMark Fire Strike | |
|---|---|---|---|
| Workload | Single compute dispatch + single draw call | Multi-pass rasterisation, tessellation, post-FX | Multi-pass rasterisation, particle physics |
| Draw calls / frame | 1 | Thousands | Thousands |
| Bottleneck (fast GPU) | CPU overhead | GPU (texture, geometry, shading) | GPU (texture, shading) |
| Memory access | Sequential SSBO read/write | Random texture fetches, render targets | Random texture fetches |
| Benefits from | Memory bandwidth, compute scheduler | Shader count, TMUs, ROPs, driver DX12 path | Shader count, TMUs, ROPs, driver DX11 path |
This difference explains why deviations exist — and why both benchmarks are needed for a complete GPU performance picture.
The 3DMark API Overhead test measures raw draw call throughput (draw calls per second) for each graphics API. This directly complements this benchmark's cross-API performance comparison by isolating driver/API overhead from GPU compute capability.
Data source:
3DMark Results/*/...-api-result.3dmark-result→ extracted viascripts/extract_3dmark.py
| GPU | Architecture | FP32 (TFLOPS) | Boost Clock | DX11-ST | DX11-MT | DX12 | Vulkan | VK / DX11 |
|---|---|---|---|---|---|---|---|---|
| RX 9070 XT | RDNA 4 (64 CU) | 48.7 † | 2970 MHz | 2.41 M | 3.05 M | 25.88 M | 39.88 M | 16.5× |
| RX 6900 XT | RDNA 2 (80 CU) | 23.04 | 2250 MHz | 2.48 M | 3.12 M | 40.78 M | 36.75 M | 14.8× |
| RX 6600 XT | RDNA 2 (32 CU) | 10.60 | 2589 MHz | 2.51 M | 3.19 M | 40.32 M | 39.25 M | 15.6× |
| Vega FE | GCN 5 (64 CU) | 13.11 | 1600 MHz | 2.22 M | 2.49 M | 26.57 M | 26.06 M | 11.8× |
| RX 580 | GCN 4 (36 CU) | 6.17 | 1340 MHz | 2.38 M | 2.30 M | 27.67 M | 26.92 M | 11.3× |
| FirePro D700 | GCN 1.0 (32 CU) | 3.48 | 850 MHz | 1.16 M | 1.03 M | 10.04 M | 9.46 M | 8.2× |
| HD 5770 | TeraScale 2 (800 SP) | 1.36 | 850 MHz | 1.60 M | 1.60 M | — | — | — |
| Ryzen 5 7600 iGPU | RDNA 2 (2 CU) | 0.56 | 2200 MHz | 2.43 M | 3.05 M | 7.22 M | 7.14 M | 2.9× |
| Adreno 640 | Adreno 6xx (768 ALU) | ~0.90 | 585 MHz | 0.16 M | 0.18 M | 0.87 M | 0.70 M | 4.3× |
† RDNA 4 uses dual-issue FP32 (each stream processor executes 2 FP32 ops/clock). Without dual-issue, the traditional calculation yields 24.3 TFLOPS — comparable methodology to the RDNA 2 and GCN numbers above.
Draw call throughput is in millions of draw calls per second. DX11-ST = single-threaded, DX11-MT = multi-threaded.
Sources: AMD official specifications, TechPowerUp GPU Database, 3DMark API Overhead test results.
-
DX11 is CPU-bound, not GPU-bound: All discrete GPUs from HD 5770 to RX 9070 XT cluster at 1.6–2.5 M DX11 single-threaded draw calls regardless of GPU power. The bottleneck is the single-threaded CPU driver path, not the GPU. DX11 multi-threading provides only 1.0–1.3× improvement on AMD (compared to NVIDIA's often 0.1× due to different driver threading models).
-
RDNA 2 has the best DX12 driver efficiency: The RX 6900 XT (40.78 M) and RX 6600 XT (40.32 M) both outperform the RX 9070 XT (25.88 M) in DX12 draw call throughput. This is likely because the RDNA 4 DX12 driver is newer and less optimised — the same pattern seen in early driver releases for previous AMD architectures.
-
RDNA 4 leads in Vulkan: The RX 9070 XT (39.88 M) achieves the highest Vulkan throughput, slightly above the RX 6600 XT (39.25 M). This aligns with this benchmark's Vulkan results where the 9070 XT performs best.
-
GCN 4/5 parity: The RX 580 (GCN 4) and Vega FE (GCN 5) show nearly identical API overhead (~27 M DX12, ~26 M Vulkan), confirming that the driver stack is the same between these two GCN generations.
-
iGPU bandwidth-limited: The Ryzen 5 7600 iGPU achieves only 7 M DX12/Vulkan despite using the same RDNA 2 driver as the 6600 XT (40 M). The 2-CU iGPU's shared DDR5 memory bandwidth (not driver overhead) limits draw call submission rate.
-
GCN 1.0 (D700) shows age: At 10 M DX12 and 9.5 M Vulkan, the 2013-era FirePro D700 achieves 25% of RDNA 2's throughput — a reasonable result given the 8-year architecture gap and older driver codepath.
-
TeraScale 2 is DX11-only: The HD 5770 has no DX12 or Vulkan support, so API overhead comparison is limited to DX11 where it achieves 1.60 M (67% of modern GPUs' DX11 rate — bottlenecked by the older CPU driver).
-
Adreno 640 (mobile SoC): At 0.16 M DX11-ST and 0.70 M Vulkan, the Snapdragon 855's GPU achieves only 7% of desktop AMD DX11 throughput and 2% of desktop Vulkan throughput. The 4.3× VK/DX11 ratio is much lower than desktop GPUs (11–17×), reflecting the mobile driver's relatively efficient DX11 path (via translation layer) and limited Vulkan command processor bandwidth.
The API Overhead results explain patterns observed in this benchmark's cross-API testing (Section 5):
-
Why Vulkan ≈ DX11 on AMD 9070 XT: The 9070 XT's Vulkan driver (39.88 M) is 16.5× faster than DX11 (2.41 M) at draw call submission. But this benchmark's single draw call per frame means the per-call overhead difference is negligible — both APIs spend < 0.01 ms on the draw call itself. The performance similarity is expected for single-dispatch workloads.
-
Why DX12 underperforms on 9070 XT: Despite DX12's explicit nature, the 9070 XT's DX12 driver (25.88 M) achieves only 65% of its Vulkan throughput. This immature DX12 driver may also explain the slightly lower DX12 FPS observed in this benchmark's cross-API comparison.
| Observation | Explanation |
|---|---|
| RDNA 4 (9070 XT) achieves 4.1× better per-CU compute than RDNA 2 (6600 XT) with 2× the CU count (64 vs 32) | Higher clocks (1.15×) + architectural improvements in scheduler, cache hierarchy, and driver codegen account for the remaining ~1.8× |
| 9070 XT outperforms 80-CU RX 6900 XT in compute (0.033 vs 0.063 ms) despite 80% CU count | RDNA 4's per-CU efficiency (~1.7× higher) is enough to overcome the 1.25× CU disadvantage |
| API ranking on 9070 XT (DX11 ≈ Vulkan > DX12 > OpenGL) differs from RTX 5090 (DX11 >> DX12 > Vulkan > OpenGL) | AMD's DX11 driver is less optimised than NVIDIA's; the gap between explicit and implicit APIs is much smaller on AMD |
| OpenGL compute overhead persists on RDNA 4 (~2.6 ms) at similar levels to RDNA 2 (~2.7 ms) | AMD's OpenGL-to-Vulkan translation layer has not improved compute dispatch overhead across generations |
| 9070 XT compute scales 59× for 16× particle increase (1M → 16M) | Super-linear: at 1M particles GPU is underutilised, per-dispatch overhead dominates; at 16M the ALUs and bandwidth are fully saturated |
| TeraScale 2 VLIW5 achieves only ~50–70% of theoretical per-SP throughput | VLIW5 slot packing inefficiency in compute shaders with irregular control flow |
| All APIs converge to 0.034 ms compute in headless mode on 9070 XT | Proves GPU compute hardware is identical across APIs; windowed differences are entirely presentation/driver overhead |
To confirm that this benchmark accurately reflects real-world GPU performance differences, results are cross-validated against 3DMark — the industry-standard graphics benchmark by UL (formerly Futuremark).
- Run this project's benchmark on each GPU (best FPS across all APIs).
- Run 3DMark Time Spy (DX12) and Fire Strike (DX11) on the same GPUs.
- Normalise all scores to a common baseline GPU (e.g. RX 580 = 1.00×).
- Compare the relative performance ratios.
If both benchmarks rank GPUs in the same order with similar ratios, it validates that this project's compute-heavy workload is a meaningful GPU performance indicator.
Baseline: RX 580 = 1.00×
| GPU | Architecture | This Benchmark | 3DMark Time Spy | 3DMark Fire Strike | Deviation (TS) | Deviation (FS) |
|---|---|---|---|---|---|---|
| RTX 5090 | Blackwell (170 SM) | 8.49× | 8.66× | 5.62× | −2.0% | +51.1% |
| RX 9070 XT | RDNA 4 (64 CU) | 1.95× | 6.49× | 4.44× | −70.0% | −56.1% |
| RX 6900 XT | RDNA 2 (80 CU) | 4.46× | 4.63× | 3.82× | −3.7% | +16.8% |
| RX 6600 XT | RDNA 2 (32 CU) | 2.01× | 2.16× | 1.92× | −7.0% | +4.9% |
| Vega FE | GCN 5 (64 CU) | 1.88× | 1.59× | 1.46× | +18.2% | +28.8% |
| RX 580 | GCN 4 (36 CU) | 1.00× | 1.00× | 1.00× | — | — |
| Zen4/5 iGPU (2 CU) | RDNA 2 | 0.36× | 0.16× | 0.15× | +119.2% | +132.0% |
| HD 5770 | TeraScale 2 | 0.21× | N/A (no DX12) | 0.10× | N/A | +106.8% |
Deviation =
(This Benchmark ratio / 3DMark ratio) − 1. Positive means our benchmark favours that GPU more than 3DMark; negative means less.
This project runs a single compute dispatch + single draw call per frame. 3DMark runs complex multi-pass rasterisation with thousands of draw calls, tessellation, post-processing, and full-screen effects. Expected differences:
| GPU | Deviation (TS) | Deviation (FS) | Reason |
|---|---|---|---|
| RTX 5090 | −2.0% | +51.1% | Time Spy near-perfect; Fire Strike deviation due to DX11 CPU overhead in 3DMark vs single-dispatch Vulkan here |
| RX 9070 XT | −70.0% | −56.1% | FPS is presentation-limited (swapchain throttling) at 1,774 FPS despite 0.033 ms compute. 3DMark exercises the full GPU; this benchmark cannot (see Section 6) |
| RX 6900 XT | −3.7% | +16.8% | Strong Time Spy correlation; Fire Strike deviation from DX11 overhead differential |
| RX 6600 XT | −7.0% | +4.9% | Good correlation across both benchmarks |
| Vega FE | +18.2% | +28.8% | HBM2's 483 GB/s bandwidth disproportionately benefits this bandwidth-bound compute workload; 3DMark's texture-heavy scenes do not leverage raw bandwidth as heavily |
| Zen4/5 iGPU (2 CU) | +119.2% | +132.0% | The iGPU's 2 CUs handle this simple compute workload efficiently (cache-friendly, low contention), but 3DMark's complex geometry/texture workloads expose the severe shader count limitation |
| HD 5770 | N/A | +106.8% | TeraScale 2 has no DX12; Fire Strike deviation similar to iGPU — simple compute overperforms relative to 3DMark's complex rasterisation |
GPUs operating in their GPU-bound regime (RTX 5090, RX 6900 XT, RX 6600 XT) show Time Spy deviations within ±10%, confirming strong correlation. The large deviations for the RX 9070 XT (presentation-limited), Vega FE (bandwidth advantage), iGPU and HD 5770 (workload mismatch) are explainable by workload characteristics and are not indicative of benchmark error.
A linear regression of project FPS vs 3DMark Time Spy scores across all GPUs yields a strong linear relationship for most GPUs, with the RX 9070 XT and iGPU as known outliers due to presentation throttling and workload mismatch respectively (see Section 6). Excluding these outliers, the remaining GPUs (RTX 5090, RX 6900 XT, RX 6600 XT, RX 580) show Time Spy deviations within ±10%, confirming the benchmark's validity for GPUs operating in their GPU-bound regime.
Charts: Run
python scripts/compare_3dmark.py --save docs/imagesto generate the normalised bar chart and correlation scatter plot (docs/images/3dmark_comparison.png,docs/images/3dmark_correlation.png).
3DMark scores are stored in scripts/3dmark_scores.json.
To auto-import from 3DMark result files:
# Import from .3dmark-result files (3DMark Advanced/Professional)
python scripts/compare_3dmark.py --import-3dmark "C:\Users\*\Documents\3DMark\*.3dmark-result"The .3dmark-result format is a ZIP archive containing arielle.xml
(benchmark scores, per-loop FPS) and si.xml (GPU name, VRAM, driver
version). The import script parses both and merges into the JSON scores file.
System Configuration
| Component | Specification |
|---|---|
| CPU | AMD Ryzen 7 9800X3D 8-Core Processor |
| Discrete GPU | NVIDIA GeForce RTX 5090 (32 GB GDDR7) |
| Integrated GPU | AMD Radeon Graphics (Zen 4, 2 CU, 2 GB shared DDR5) |
| OS | Windows 11 25H2 |
| Resolution | 1280 × 720 |
| V-Sync | OFF |
| Memory Mode | Device-local (staging buffer → VRAM on dGPU) |
| Metric | Vulkan | DirectX 12 | DirectX 11 | OpenGL 4.3 |
|---|---|---|---|---|
| Avg FPS | 3,611 | 6,547 | 8,955 | 2,442 |
| Avg GPU Time | 0.094 ms | 0.065 ms | 0.104 ms | 0.087 ms |
| Avg Frame Time | 0.277 ms | 0.153 ms | 0.112 ms | 0.409 ms |
| CPU Overhead / Frame | 0.183 ms | 0.088 ms | 0.008 ms | 0.322 ms |
| GPU Utilisation | 33.9% | 42.4% | 93.2% | 21.2% |
| Bottleneck | CPU-bound | CPU-bound | GPU-bound | CPU-bound |
All four APIs deliver nearly identical GPU execution times (0.065–0.104 ms), confirming the GPU-side workload is equivalent. The FPS difference is entirely driven by per-frame CPU overhead.
Ranking: DX11 > DX12 > Vulkan > OpenGL
Why DX11 is fastest here:
- DX11 is an implicit API — the driver handles command batching, resource state tracking, and barrier insertion internally. NVIDIA's DX11 driver path has been optimised for over a decade, making it extremely efficient for simple, single-threaded workloads.
- Per-frame CPU overhead is only 0.008 ms, leaving the GPU as the actual bottleneck (93.2% utilisation).
Why DX12/Vulkan are slower here:
- Both are explicit APIs requiring the application to manually manage command allocators, fences, resource barriers, and descriptor heaps/sets.
- This shifts work from the driver to application code, adding 10–20× more CPU overhead per frame compared to DX11.
- With only 1 compute dispatch + 1 draw call per frame, there is no opportunity for multi-threaded command recording — the very feature that justifies explicit APIs in complex scenes.
Why OpenGL is slowest here:
- OpenGL has the highest per-frame CPU overhead at 0.322 ms — roughly 40× more than DX11.
glfwSwapBufferson Windows goes through WGL, which has less efficient frame queue management than DXGI'sPresentpath.- OpenGL's global state machine model means every
glUseProgram,glBindBuffer, andglBindVertexArraycall triggers internal driver state validation, accumulating significant overhead even with minimal draw calls. - Despite being an implicit API like DX11, OpenGL's Windows driver path has received far less optimisation from NVIDIA in recent years, as industry focus has shifted to Vulkan and DirectX.
- However, OpenGL still produces the second-fastest GPU execution time (0.087 ms), confirming the bottleneck is purely in the CPU-side driver, not in the shader or buffer management.
When DX12/Vulkan win:
Explicit APIs excel when a scene contains hundreds or thousands of draw calls. In that scenario, DX11's single-threaded driver becomes the bottleneck, while DX12/Vulkan can parallelise command recording across multiple CPU threads, reducing total CPU time proportionally.
When OpenGL makes sense:
OpenGL 4.3 remains the most portable option — it runs on Windows, Linux, and macOS (legacy profile) without requiring Vulkan drivers or platform-specific APIs. For workloads that are GPU-bound (high particle counts, complex shaders), OpenGL's higher CPU overhead becomes negligible relative to total frame time.
AMD comparison: The RX 9070 XT (RDNA 4) shows a very different API ranking: DX11 ≈ Vulkan (1,774 / 1,751 FPS) > DX12 (1,609 FPS) >> OpenGL (253 FPS). The DX11-over-Vulkan advantage shrinks from 2.5× on NVIDIA to 1.01× on AMD, reflecting AMD's less optimised DX11 driver. See Section 5 for the full RX 9070 XT cross-API analysis.
| Metric | RTX 5090 (Discrete) | RX 9070 XT (Discrete) | AMD Radeon iGPU (Integrated) |
|---|---|---|---|
| Avg FPS | 2,700+ | 1,751 | ~320 |
| Compute | 0.035 ms | 0.033 ms | 1.47 ms |
| Render | 0.045 ms | 0.408 ms | 1.5 ms |
| Total GPU | 0.08 ms | 0.446 ms | ~3.0 ms |
| Ratio | 1× | ~5.6× slower | ~37× slower |
The RTX 5090 (21,760 CUDA cores, ~3,000 GB/s bandwidth) outperforms the Zen 4 iGPU (128 shaders, ~50 GB/s shared DDR5) by approximately 37× in GPU execution time. This aligns with the memory bandwidth ratio (~60×), confirming the particle simulation is bandwidth-bound rather than compute-bound at this scale.
The RX 9070 XT sits between these extremes: its compute time (0.033 ms) is comparable to the RTX 5090 (0.035 ms), but its total GPU time (0.446 ms) is 5.6× higher due to swapchain semaphore wait pollution inflating the render timestamp (see Section 6). In headless mode, the 9070 XT achieves 21,260 FPS — only ~15% behind the RTX 5090's headless throughput.
| # | API | Avg FPS | Compute (ms) | Render (ms) | Total GPU (ms) |
|---|---|---|---|---|---|
| 1 | DX12 | 611.7 | 0.581 | 0.730 | 1.312 |
| 2 | Vulkan | 539.9 | 0.628 | 1.006 | 1.635 |
| 3 | DX11 | 470.0 | 0.605 | 0.771 | 2.000 |
| 4 | OpenGL | 457.6 | 0.580 | 0.969 | 1.550 |
At 16M particles the RTX 5090 remains fast enough that DX12 retakes the lead from DX11 (612 vs 470 FPS). The workload is now GPU-bound, so DX11's low CPU overhead advantage disappears and its higher total GPU time (2.000 ms, likely due to implicit barrier overhead) becomes the bottleneck.
| # | API | Avg FPS | Compute (ms) | Total GPU (ms) |
|---|---|---|---|---|
| 1 | DX11 | 37,564 | 0.000 | 0.000 |
| 2 | OpenGL | 35,110 | 0.018 | 0.020 |
| 3 | Vulkan | 26,358 | 0.023 | 0.025 |
| 4 | DX12 | 24,758 | 0.014 | 0.014 |
DX11 timestamp anomaly: The RTX 5090's DX11 headless mode reports 0.000 ms for all GPU timing metrics, suggesting NVIDIA's DX11 driver does not support timestamp queries in headless/compute-only mode. The 37,564 FPS figure is valid (derived from CPU-side frame timing), but the GPU time breakdown is unavailable.
Cross-vendor headless comparison (1M particles, best API):
| GPU | Best Headless FPS | Best API | Price (MSRP) | FPS per $ |
|---|---|---|---|---|
| RTX 5090 | 37,564 | DX11 | $1,999 | 18.8 |
| RX 9070 XT | 21,354 | DX12 | $599 | 35.6 |
| RX 6900 XT | 15,950 | DX12 | $999 (launched) | 16.0 |
The RTX 5090 achieves 1.76× the headless throughput of the RX 9070 XT — a significant lead, but far from the ~2.2× TFLOPS ratio (105 vs 48.7 TFLOPS) or the ~2.8× bandwidth ratio (1,792 vs 640 GB/s). On a price-performance basis, the RX 9070 XT delivers 1.9× the FPS per dollar of the RTX 5090 for this compute workload. Even comparing Vulkan-to-Vulkan (26,358 vs 21,260 FPS), the 5090's lead narrows to just 1.24×.
This confirms that for bandwidth-bound compute workloads, mid-range GPUs offer substantially better value than flagships — the RTX 5090's additional CUDA cores and memory bandwidth face diminishing returns when the workload cannot saturate them.
| Memory Mode | Compute | Render | Total GPU | FPS |
|---|---|---|---|---|
| Device-local (default) | 0.035 ms | 0.045 ms | 0.08 ms | 2,700+ |
Host-visible (--host-memory) |
1.25 ms | 0.15 ms | 1.4 ms | ~600 |
Using host-visible memory (system RAM accessed over PCIe) instead of device-local VRAM causes a 35× increase in compute time on a discrete GPU. The compute shader reads/writes particle data every frame — over PCIe, this becomes the dominant bottleneck.
On an integrated GPU, this penalty disappears because host-visible and device-local memory both reside in the same physical DDR5, making the distinction meaningless.
WARP (Windows Advanced Rasterisation Platform) is Microsoft's CPU-based software rasteriser bundled with every modern Windows installation. It runs the entire graphics pipeline on the CPU using SIMD (SSE/AVX) and multi-threading, serving as both a correctness reference and a fallback when no hardware GPU driver is available.
Native API support: WARP natively implements Direct3D 11 and Direct3D 12 only. It does not implement Vulkan or OpenGL. If Vulkan is reported as available on a WARP-only system, this is provided by Mesa Dozen — a Vulkan-on-D3D12 translation layer distributed via the Microsoft Store's OpenCL, OpenGL & Vulkan Compatibility Pack. Similarly, OpenGL support on WARP comes from OpenGLOn12 in the same compatibility pack. Both layers translate their respective API calls to D3D12, which WARP then executes on the CPU.
| Metric | RTX 5090 / DX12 (Hardware) | WARP / DX12 (Software) |
|---|---|---|
| Avg FPS | 6,547 | 83 |
| Compute | 0.014 ms | 1.1 ms |
| Render | 0.050 ms | 10.6 ms |
| Total | 0.065 ms | 11.7 ms |
WARP demonstrates a ~80× performance gap compared to hardware GPU execution, which is expected for CPU-based software rasterisation.
| Metric | WARP + DX11 | WARP + DX12 |
|---|---|---|
| Avg FPS | 52 | 83 |
| Timestamp queries | Not available | 11.7 ms total |
On hardware GPUs, DX11 outperforms DX12/Vulkan because the driver's implicit state management is highly optimised and adds negligible overhead. On WARP, the result reverses: DX12 is 60% faster than DX11.
Why the reversal:
- DX11's driver layer becomes pure overhead. On a hardware GPU, the DX11 runtime performs implicit resource state tracking, dependency analysis, and barrier insertion to optimise GPU command submission. When the "GPU" is WARP (a CPU-based software renderer), there is no hardware to optimise for — this entire layer is wasted CPU work.
- DX12's thin runtime is a better fit. DX12's explicit model has minimal runtime between the application and the execution engine. The application specifies exactly what to do, and WARP executes it directly with less translation overhead.
- WARP's DX12 implementation is more modern. DX12 (introduced 2015) benefits from a newer WARP backend that may leverage more efficient internal scheduling compared to the legacy DX11 WARP path.
This observation reinforces that the DX11 driver's "free optimisation" is specifically valuable for hardware GPU command submission — when that hardware is absent, the optimisation layer becomes a liability.
On systems without a native Vulkan ICD (e.g. Windows on ARM VMs, virtual GPUs), the Vulkan loader may enumerate "Microsoft Basic Render Driver" as a Vulkan physical device. The full call chain is:
Vulkan application → Dozen (Vulkan → D3D12) → WARP (D3D12 → CPU)
Dozen is distributed as part of the OpenCL, OpenGL & Vulkan Compatibility Pack from the Microsoft Store (D3DMappingLayers app package). Windows 11 may install this pack automatically on devices that lack native Vulkan/OpenGL drivers, particularly Windows on ARM devices and virtual machines.
On systems with a hardware Vulkan ICD (e.g. NVIDIA, AMD), the Dozen/WARP device is typically not enumerated or is deprioritised by the Vulkan loader. Selecting Vulkan on a WARP-only system is functionally identical to selecting DX12 on WARP — both end up as CPU-based software rendering, with Dozen adding a thin additional translation layer.
| Component | Specification |
|---|---|
| CPU | AMD Ryzen 5 7600 6-Core Processor |
| Discrete GPU | AMD Radeon HD 5770 (757 MB GDDR5, TeraScale 2, 2009) |
| Integrated GPU | AMD Radeon Graphics (Zen 4 / RDNA 2, 2 CU, shared DDR5) |
| OS | Windows 11 (NT 10.0.26200) |
| Resolution | 1280 × 720 |
| V-Sync | OFF |
| Memory Mode | Device-local |
The HD 5770 only supports DX11 and OpenGL 4.3 — no Vulkan or DX12 drivers exist for TeraScale 2 hardware. The Ryzen 7600 iGPU supports all four APIs.
| # | API | GPU | Avg FPS | Compute (ms) | Render (ms) | Total GPU (ms) | Utilisation |
|---|---|---|---|---|---|---|---|
| 1 | DX12 | Radeon iGPU (RDNA 2) | 313 | — | — | 2.956 | — |
| 2 | Vulkan | Radeon iGPU (RDNA 2) | 275 | — | — | 3.426 | — |
| 3 | OpenGL | Radeon iGPU (RDNA 2) | 275 | — | — | 3.391 | — |
| 4 | OpenGL | HD 5770 (TeraScale 2) | 193 | 1.789 | 3.025 | 4.820 | 93.1% |
| 5 | DX11 | Radeon iGPU (RDNA 2) | 190 | — | — | 5.017 | — |
| 6 | DX11 | HD 5770 (TeraScale 2) | 111 | 1.078 | 2.715 | 8.973 | 99.6% |
| 7 | DX12 | WARP (CPU) | 64 | — | — | 15.234 | — |
| 8 | DX11 | WARP (CPU) | 45 | — | — | 21.954 | — |
| API | HD 5770 FPS | iGPU FPS | iGPU Advantage |
|---|---|---|---|
| DX11 | 111 | 190 | +71% |
| OpenGL | 193 | 275 | +42% |
The RDNA 2 integrated GPU outperforms the HD 5770 discrete GPU in every comparable API, despite having far fewer hardware resources on paper.
| HD 5770 | Ryzen 5 7600 iGPU | |
|---|---|---|
| Architecture | TeraScale 2 (2009) | RDNA 2 (2022) |
| Stream Processors | 800 (160 × VLIW5) | 128 (2 CU × 64) |
| Core Clock | 850 MHz | 2,200 MHz |
| FP32 TFLOPS | ~1.36 | ~0.56 |
| Memory | 1 GB GDDR5, ~76.8 GB/s | Shared DDR5, ~83 GB/s |
The HD 5770 has 2.4× more raw FP32 TFLOPS than the iGPU, yet it is 42–71% slower in this compute benchmark. This inversion demonstrates that TFLOPS alone is a poor predictor of real-world compute shader performance.
This is not anomalous — TFLOPS comparisons across different architectures are unreliable as an industry rule of thumb. Well-documented examples include AMD Vega 64 (13.7 TFLOPS) losing to NVIDIA GTX 1080 (8.9 TFLOPS) in many gaming and compute workloads, and Intel Arc A770 (19.7 TFLOPS) underperforming against the RTX 3060 (12.7 TFLOPS) at launch despite a 55% TFLOPS advantage. TFLOPS measures only the theoretical rate of fused multiply-add operations — it says nothing about whether the ALUs can actually be kept fed with data and useful instructions. Cache hit rates, memory bandwidth, scheduling efficiency, VLIW slot utilisation, and driver code generation quality all determine how much of the theoretical peak is realised in practice.
The discrepancy between TFLOPS rankings and benchmark results is itself a validation of the test. If results tracked TFLOPS perfectly, it would suggest the benchmark is merely saturating ALU throughput with a trivially parallel workload — essentially an artificial peak-FLOPS test. The fact that a 0.56 TFLOPS GPU outperforms a 1.36 TFLOPS GPU confirms that this benchmark exercises real-world bottlenecks — memory access patterns, compute scheduler overhead, wave occupancy, and driver-side code generation — rather than measuring a synthetic upper bound.
Architecture efficiency matters more than shader count. TeraScale 2 uses a VLIW5 (Very Long Instruction Word) design where each "stream processor" is actually five tightly coupled ALUs that must execute in lockstep. If the compiler cannot fill all five slots (a common occurrence for compute shaders with irregular control flow), the vacant slots are wasted. Real-world VLIW5 utilisation in compute workloads is estimated at 50–70%, reducing the HD 5770's effective throughput to roughly 0.7–0.95 TFLOPS.
RDNA 2, by contrast, uses a scalar + SIMD32 design where each compute unit contains two independent SIMD32 units. Every lane executes useful work on every clock — there is no VLIW packing problem. At 2,200 MHz, the 128 shaders deliver nearly their full 0.56 TFLOPS.
The real gap is far smaller than the spec sheet suggests. After accounting for VLIW5 utilisation losses, the effective compute advantage shrinks from the theoretical 2.4× (1.36 vs 0.56 TFLOPS) down to roughly 1.3–1.7× (0.7–0.95 vs 0.56 TFLOPS). The remaining factors below — memory bandwidth parity, driver quality, and compute scheduler maturity — are more than sufficient to close this residual gap and tip the balance in the iGPU's favour.
Compute shader support maturity. The HD 5770 was designed primarily for DirectX 11-era pixel and vertex shading. Its compute shader support (DirectCompute 5.0) was a first-generation implementation with limited occupancy, no asynchronous compute queues, and restricted shared memory bandwidth. RDNA 2 treats compute as a first-class workload with dedicated hardware schedulers, LDS (Local Data Share) bandwidth matched to ALU throughput, and fine-grained wave management.
Driver optimisation. AMD's current Radeon Software Adrenalin Edition drivers for RDNA 2 are actively maintained and optimised. The HD 5770's legacy Crimson Edition drivers (version 16.2.1, Mar 2016) have not received performance updates in over a decade (actually, it‘s real 10 years, now it's Mar 2026). Compute shader code generation for TeraScale 2 was never a priority — these drivers were written when GPU compute was still in its infancy.
Memory bandwidth parity. The HD 5770's theoretical advantage in dedicated GDDR5 is largely neutralised here. Its 76.8 GB/s bandwidth is slightly below the iGPU's ~83 GB/s from dual-channel DDR5-6000 C28. For a bandwidth-sensitive particle simulation, this effectively levels the playing field — or tilts it slightly in the iGPU's favour.
Very likely yes, for traditional 3D rendering workloads. The HD 5770 has 6.25× more shader units, 5× more texture mapping units, and 4× more render output units than the 2-CU iGPU. In a conventional rasterisation pipeline — vertex processing, texture sampling, pixel shading, and blending — these fixed-function resources matter far more than per-CU compute efficiency.
Online gaming benchmarks broadly confirm this: the HD 5770 can run older titles (pre-2015) at low-medium settings, whereas the Ryzen 7600 iGPU struggles to maintain playable frame rates in the same scenarios.
The key insight: This benchmark is a compute-first workload — a particle simulation driven by a compute shader, with a simple instanced rendering pass for visualisation. It exercises the GPU's general-purpose compute pipeline, not its fixed-function rasterisation hardware. The result is a measure of compute shader throughput and scheduling efficiency, where architectural modernity dominates raw shader count.
This makes the benchmark a useful complement to traditional GPU tests. A gaming benchmark tells you how fast a GPU can rasterise triangles; this benchmark tells you how efficiently it can execute general-purpose parallel computation — a workload increasingly relevant to physics simulation, machine learning inference, post-processing, and scientific computing.
Unlike Vulkan, DirectX 11, and DirectX 12, OpenGL has no standard API for enumerating or selecting a specific GPU on a multi-GPU system. Each of the other backends provides an adapter/device enumeration mechanism:
| API | GPU Enumeration | Per-GPU Selection |
|---|---|---|
| Vulkan | vkEnumeratePhysicalDevices |
Create device on any enumerated physical device |
| DirectX 12 | IDXGIFactory::EnumAdapters |
Pass chosen adapter to D3D12CreateDevice |
| DirectX 11 | IDXGIFactory::EnumAdapters |
Pass chosen adapter to D3D11CreateDevice |
| OpenGL | None | OS/driver decides |
On Windows, the OpenGL context is created by the OS display driver model (WDDM), which assigns the GPU based on system-level configuration. The application has no standard API to override this at runtime.
Available workarounds (limited):
| Method | Scope | Limitation |
|---|---|---|
NvOptimusEnablement export symbol |
Forces discrete NVIDIA GPU on Optimus laptops | Only works on NVIDIA + Intel hybrid laptops; no effect on desktop multi-GPU |
AmdPowerXpressRequestHighPerformance export symbol |
Forces discrete AMD GPU on switchable graphics laptops | Same — laptop-only, binary choice (discrete vs integrated) |
WGL_NV_gpu_affinity extension |
Per-GPU context creation | Quadro professional cards only — not available on GeForce/consumer GPUs |
| Windows Graphics Settings panel | Per-executable GPU assignment | Requires manual user configuration outside the application |
On the test system (RTX 5090 + AMD Radeon iGPU desktop), none of the programmatic methods are effective — the only way to force OpenGL onto the integrated GPU is through the Windows Graphics Settings panel.
Linux provides significantly better OpenGL GPU selection:
| Method | Scope | How |
|---|---|---|
DRI_PRIME=N environment variable |
Per-process GPU selection (Mesa drivers) | DRI_PRIME=1 ./gpu_benchmark |
__NV_PRIME_RENDER_OFFLOAD=1 |
Per-process offload to NVIDIA GPU | __NV_PRIME_RENDER_OFFLOAD=1 __GLX_VENDOR_LIBRARY_NAME=nvidia ./gpu_benchmark |
EGL_EXT_platform_device |
Programmatic per-GPU EGLDisplay creation | Requires EGL instead of GLX; Mesa 23.3+ |
EGL_EXT_explicit_device |
Same, with native windowing support | Mesa 23.3+ |
The application detects Linux at runtime and uses DRI_PRIME to route OpenGL to the user's requested GPU index.
This limitation means OpenGL cross-GPU comparisons on Windows require manual configuration, whereas all other backends support interactive GPU selection within the application. On Linux, DRI_PRIME provides equivalent functionality to other backends' built-in GPU selection.
DX11 is the only API in the benchmark where GPU timestamp queries can silently fail to produce results. Vulkan and DX12 always return timestamp values regardless of GPU clock state. DX11, by contrast, uses a D3D11_QUERY_TIMESTAMP_DISJOINT wrapper that can actively refuse to return data.
Three distinct failure modes were observed during testing:
Affected: Windows on ARM virtual machines (SVGA virtual GPU driver).
ID3D11Device::CreateQuery succeeds for both D3D11_QUERY_TIMESTAMP and D3D11_QUERY_TIMESTAMP_DISJOINT, and the application reports timestamps as "enabled". However, ID3D11DeviceContext::GetData for the disjoint query perpetually returns S_FALSE — the result is never ready.
This is a driver limitation: the virtual GPU driver accepts query creation but does not implement the hardware counters needed to resolve them. No application-level workaround exists.
See docs/woa-dx11-timestamp-issue.md for a detailed write-up.
Affected: Integrated GPUs under fluctuating load, discrete GPUs during power-state transitions.
The D3D11_QUERY_DATA_TIMESTAMP_DISJOINT structure contains a Disjoint boolean. When TRUE, it signals that the GPU's clock frequency changed during the frame (P-state transition, thermal throttling, power-saving downclock), making the timestamp-to-millisecond conversion unreliable.
The D3D11 specification recommends discarding the entire frame's timing data when Disjoint = TRUE. If the GPU is frequently switching power states — common on integrated GPUs under variable load, or during the first few seconds of a benchmark run while the GPU ramps up — this can result in many consecutive frames with no timing data.
This is not a driver bug. It is a deliberate DX11 design choice to prioritise timestamp accuracy over availability.
Key difference from Vulkan/DX12: Neither Vulkan nor DX12 has a Disjoint concept. Their timestamp queries always return values based on a fixed timestampPeriod / Frequency, even if the GPU clock changes mid-frame. The precision may degrade slightly, but data is never withheld entirely. This is why Vulkan and DX12 report timestamps reliably in scenarios where DX11 reports none.
Mitigation implemented: The application now caches the last known stable frequency (lastGoodFrequency). When Disjoint = TRUE, timestamps are still read and converted using the cached frequency rather than being discarded. This mirrors the behaviour of Vulkan/DX12 — accepting marginally less precise data in exchange for continuous availability.
Affected: Slow GPUs (integrated, software renderer) under high particle counts.
If the GPU takes significantly longer than one frame to process submitted work, the application may attempt to read a query result before the GPU has finished writing it. GetData returns S_FALSE because the query genuinely hasn't resolved yet — not because the driver doesn't support it.
Fix: The ring buffer was increased from 4 to 8 slots, and GetData retries were increased to 128 with periodic Sleep(1) yields, giving slow GPUs more time to resolve queries.
| Scenario | Root Cause | Driver Bug? | Fix |
|---|---|---|---|
| WoA virtual GPU — never returns data | Driver doesn't implement timestamp counters | Yes | None (graceful fallback to CPU-only timing) |
| iGPU / dGPU ramp-up — intermittent gaps | Disjoint = TRUE during clock transitions |
No (spec behaviour) | Use cached frequency instead of discarding |
| Slow GPU — first N frames missing | Query not resolved before read | No (pipeline depth) | Deeper ring buffer (8 slots) + retry with Sleep |
| WARP DX11 — works after warm-up | Combination of 6b and 6c | No | Same mitigations as above |
Each backend uses its own API's timestamp mechanism — there is no cross-API data sharing. The same GPU executing the same workload produces nearly identical execution times (0.065–0.104 ms on RTX 5090), but the measurement infrastructure differs significantly:
| Vulkan | DX12 | DX11 | OpenGL | |
|---|---|---|---|---|
| Write | vkCmdWriteTimestamp |
EndQuery → ID3D12QueryHeap |
context->End(query) |
glQueryCounter(GL_TIMESTAMP) |
| Read | vkGetQueryPoolResults with WAIT_BIT |
ResolveQueryData → readback buffer |
GetData (CPU polling) |
glGetQueryObjectui64v |
| Synchronisation | GPU-side wait (guaranteed ready) | GPU-side resolve (ordered in command list) | CPU polls until S_OK (may never arrive) |
CPU polls (typically resolves quickly) |
| Disjoint / clock check | None | None | Required (D3D11_QUERY_TIMESTAMP_DISJOINT) |
None |
| Counter frequency | Fixed (timestampPeriod), independent of core clock |
Fixed (GetTimestampFrequency), independent of core clock |
May vary with GPU core clock | Fixed, monotonic counter |
| Clock-change handling | Returns data; timer may reset across submissions† | Returns data; stable-clock design, no resets | Refuses data if Disjoint = TRUE |
Returns data, counter is monotonic |
| First-frame data | Yes | Yes | No (ring buffer warm-up required) | Yes (after 1–2 frame delay) |
Vulkan caveat: The Vulkan spec notes that power management events (e.g. GPU idle → active transitions) can reset the timestamp counter on some implementations. This affects cross-submission comparisons only — timestamps within the same command buffer are always reliably comparable. The VK_EXT_calibrated_timestamps extension provides monotonic timestamps immune to power events, but is not required for within-frame profiling. In this benchmark, all four timestamps (compute begin/end, render begin/end) are recorded within a single command buffer, so power-state resets do not affect the results.
DX12's stable-clock design: DX12 goes further than Vulkan by explicitly stabilising the GPU clock for timestamp purposes. Two timestamps within the same command list are always comparable, and two timestamps from different command lists are also reliable as long as the GPU did not idle between them. There is no Disjoint equivalent — the API guarantees clock stability by design.
DX11 is the only API that can actively withhold timestamp data based on GPU clock stability. DX12 and Vulkan both use a fixed counter frequency independent of the GPU core clock, so frequency scaling and P-state transitions do not invalidate their results. This makes DX11 the most fragile timestamp implementation from an application developer's perspective, despite the underlying GPU hardware being identical across all backends.
OpenGL compute shader performance on AMD GPUs is significantly lower than Vulkan / DX12 / DX11, with older architectures affected most severely.
| GPU | Architecture | OpenGL Compute ms | Vulkan Compute ms | Ratio |
|---|---|---|---|---|
| RTX 5090 (reference) | Blackwell | 0.019 | 0.019 | 1.0× |
| Radeon Graphics (iGPU) | RDNA 2 | 1.489 | 0.758 | 2.0× |
| RX 9070 XT | RDNA 4 | 2.612 | 0.033 | 79.2× |
| RX 6900 XT | RDNA 2 | 2.742 | 0.184 | 14.9× |
| RX 6600 XT | RDNA 2 | 2.719 | 0.240 | 11.3× |
| Vega Frontier Edition | Vega (GCN 5) | 3.321 | 0.368 | 9.0× |
| RX 580 | Polaris (GCN 4) | 18.913 | 0.362 | 52.3× |
On NVIDIA, OpenGL and Vulkan compute times are nearly identical. On AMD, OpenGL compute is 9–52× slower depending on architecture generation.
| GPU | OpenGL FPS | Vulkan FPS | DX11 FPS |
|---|---|---|---|
| RX 9070 XT | 253 | 1,751 | 1,774 |
| RX 6900 XT | 229 | 2866 | 4107 |
| RX 6600 XT | 180 | — | — |
| RX 580 | 42 | 783 | 755 |
The RX 580's OpenGL score (42 FPS) is lower than the Ryzen 5 7600 CPU-based WARP software renderer running DX11 (44–53 FPS).
This is a well-documented AMD Windows OpenGL driver limitation, not a code issue:
-
Same code, different results: The identical OpenGL compute path achieves 2062 FPS on RTX 5090 (GPU time 0.019 ms), confirming the shader and API usage are correct.
-
Observed per-dispatch overhead: On all GCN/RDNA GPUs tested in this benchmark, OpenGL
glDispatchComputeexhibits a consistent overhead of ~2.7 ms (RDNA 2) to ~18.9 ms (GCN 4), independent of GPU compute capability — the RX 6900 XT and RX 6600 XT show nearly identical compute times despite having 80 vs 32 CUs. While AMD's general OpenGL performance issues are well-documented (see below), specific quantification of per-dispatch compute overhead does not appear to have been published elsewhere — this benchmark may be the first to isolate and measure it. This is demonstrated by the following comparison:Metric HD 5770 OpenGL RX 6600 XT OpenGL RX 6600 XT Vulkan Compute 1.794 ms 2.719 ms 0.270 ms Render 3.018 ms 2.322 ms 0.379 ms Total GPU 4.818 ms 5.148 ms 0.649 ms FPS 188 180 1239 The RX 6600 XT's Vulkan compute time (0.270 ms) proves the hardware is 10× faster than what OpenGL reports (2.719 ms). The ~2.7 ms figure appears to be a driver-level overhead floor observed consistently across all GCN/RDNA hardware tested, rather than a reflection of GPU capability. The HD 5770 — a vastly weaker GPU from 2009 — achieves lower OpenGL compute times (1.794 ms) and higher FPS (188 vs 180) than the RX 6600 XT simply because it uses a completely different TeraScale driver stack that does not exhibit this overhead.
-
Known industry issue: Multiple major projects and community reports have documented AMD's OpenGL performance gap on Windows:
- RPCS3 (PS3 emulator) filed issue #11197 "radeon: Poor state of Windows OpenGL drivers", describing AMD's OpenGL performance as "disastrous."
- PCSX2 (PS2 emulator) documented the problem in their wiki: "OpenGL and AMD GPUs - All you need to know", noting OpenGL runs 10–70% slower compared to Direct3D on AMD GPUs.
- The Khronos Community Forums contain threads reporting execution differences between NVIDIA and AMD GPUs in compute shaders, with RX 580 and RX 5700 exhibiting issues during
dispatchCompute. - Tom's Hardware Forum and AMD Community have numerous user reports of poor AMD OpenGL performance.
- AMD acknowledged the issue and rewrote their OpenGL driver to internally translate calls to Vulkan (Adrenalin 22.7.1 / WDDM 3.1), achieving up to 55% improvement in Unigine Valley and 79% in Minecraft. VideoCardz reported up to 92% improvement in Minecraft Java Edition, and independent benchmarks by Nemez and OC3D confirmed AMD even surpassed NVIDIA in some Minecraft scenarios after the driver update.
- However, the translation layer primarily optimises the rendering path (draw calls), not compute dispatch. The applications that saw 79–92% improvements — Minecraft, Unigine — are rendering-bound workloads dominated by vertex/fragment shaders and draw calls. Our benchmark's bottleneck is in
glDispatchCompute, a relatively niche OpenGL usage pattern. This explains why even on the latest AMD drivers (Adrenalin 26.3.1, tested on the RX 6600 XT), the OpenGL compute overhead floor of ~2.7 ms persists — the translation layer offers limited benefit for this code path. Applications requiring GPU compute have largely migrated to Vulkan, DX12, or CUDA, reducing the incentive for AMD to optimise OpenGL compute dispatch specifically.
-
GCN/Polaris most affected: Since Adrenalin 23.9, AMD moved GCN (Polaris/Vega) to a maintenance-only driver branch with no new performance optimisations. The OpenGL-to-Vulkan translation layer improvements may not have been fully applied to these legacy architectures, explaining the extreme 52× gap on the RX 580.
-
TeraScale counterexample: The HD 5770 (TeraScale 2 / Evergreen) uses a completely different, legacy driver stack and does not exhibit this OpenGL compute overhead. As shown in the table above, it outperforms GCN/RDNA 2 cards in OpenGL despite being vastly inferior hardware. According to Chips and Cheese's architectural analysis, GCN completely rewrote AMD's GPU architecture and driver stack from TeraScale's VLIW design to a scalar SIMD model — meaning the OpenGL driver codebases are entirely separate. The overhead problem was introduced in the GCN/RDNA driver branch, not inherited from TeraScale. Notably, the HD 5770 shows the opposite pattern on DX11: compute + render = 3.87 ms, but total GPU time = 9.15 ms — a 5.3 ms synchronisation overhead between the compute and render stages, likely due to immature compute shader support on this early DX11-era architecture.
The OpenGL compute path in this benchmark is already minimal — one glDispatchCompute call and one glMemoryBarrier per frame. There is no room to reduce dispatch frequency, batch operations, or eliminate synchronisation. Alternative approaches such as glDispatchComputeIndirect or persistent mapped buffers do not address the bottleneck, as the overhead originates within the driver's internal dispatch path, not in data transfer or API call volume.
This suggests that when the performance bottleneck is a driver-level fixed cost, application-level optimisation cannot break through the ceiling. The only effective solution is to use an API that avoids this overhead entirely — Vulkan achieves 0.270 ms for the same compute workload that takes 2.719 ms through OpenGL on the same hardware (RX 6600 XT), a 10× improvement with identical shader logic.
AMD's OpenGL-to-Vulkan translation layer excels at optimising high-volume rendering workloads. Minecraft Java Edition issues thousands of draw calls per frame (blocks, entities, particles, UI), each carrying CPU-side overhead for state changes and submission. The translation layer batches these into efficient Vulkan command buffers, dramatically reducing the per-call cost — e.g., 1000 draw calls × 0.1 ms overhead each = 100 ms, batched down to a few grouped submissions at ~10 ms total.
This benchmark has the opposite profile: one single glDispatchCompute call per frame with a ~2.7 ms observed driver overhead. The translation layer's batching strategy cannot help here — there is nothing to batch. The overhead appears to be a per-dispatch fixed cost within the driver, not an accumulation of many small costs that can be amortised. This is why Minecraft saw up to 92% improvement while our compute workload on the same latest drivers (Adrenalin 26.3.1) shows no meaningful change.
Note: The Khronos Community Forums contain a thread on
glDispatchComputecalling overhead, but that discussion reports ~0.2 ms overhead on older NVIDIA GPUs (GTX 560/470), not AMD. The ~2.7 ms overhead observed in this benchmark on AMD GCN/RDNA hardware is an order of magnitude larger and does not appear to have been specifically documented elsewhere.
The GTX 970 exhibits a striking reversal in API performance ranking compared to the RTX 5090:
| API | Compute | Render | Compute+Render | Total GPU | FPS | Bottleneck |
|---|---|---|---|---|---|---|
| Vulkan | 0.434 ms | 0.663 ms | 1.097 ms | 1.098 ms | 718.8 | Balanced |
| OpenGL | 0.431 ms | 0.791 ms | 1.222 ms | 1.226 ms | 642.2 | Balanced |
| DX12 | 0.443 ms | 0.535 ms | 0.978 ms | 0.977 ms | 291.1 | CPU-bound |
| DX11 | 0.435 ms | 0.873 ms | 1.308 ms | 3.355 ms | 280.3 | GPU-bound |
On the RTX 5090, DX11 is the fastest API (7736 FPS). On the GTX 970, it is the slowest (280 FPS). The cause is visible in the numbers:
- DX11: Compute + render sum to only 1.308 ms, but total GPU time is 3.355 ms — a ~2 ms synchronisation overhead between the compute and render stages. Maxwell's DX11 driver appears to insert a costly pipeline flush/barrier when transitioning from compute dispatch to draw calls. The RTX 5090 (Blackwell) shows no such overhead (compute + render ≈ total GPU time).
- DX12: GPU time is actually the fastest (0.977 ms), but FPS is only 291.1 — a massive 2.5 ms CPU overhead (frame time 3.435 ms − GPU 0.977 ms). This reflects the high CPU-side cost of DX12 command recording on an older driver/architecture combination.
- Vulkan: Best overall balance — GPU time 1.098 ms with only 0.293 ms CPU overhead. The explicit API model with pre-recorded command buffers works well even on older hardware.
- OpenGL: NVIDIA's OpenGL driver performs well (unlike AMD), with GPU time 1.226 ms and minimal CPU overhead.
This demonstrates that API performance rankings are not universal — they depend on GPU architecture and driver maturity. DX11's implicit driver model excels on modern NVIDIA hardware (where the driver has been refined over a decade) but introduces overhead on older architectures where compute–render transitions were not as optimised. The RX 9070 XT (RDNA 4) shows yet another pattern: DX11 ≈ Vulkan ≈ DX12, with only OpenGL significantly behind — AMD's DX11 advantage over explicit APIs is negligible compared to NVIDIA's.
The earlier version of this report incorrectly concluded that Feature Level 10_0
could not execute compute shaders at all. The observed failure was real, but its
cause was in this benchmark: the DX11 backend created an FL10_0 device and then
unconditionally compiled cs_5_0, vs_5_0, and ps_5_0. Naturally,
CreateComputeShader rejected that Shader Model 5 bytecode on an SM4 device.
Direct3D 11 exposes an optional, genuine DirectCompute 4.x path for Direct3D
10.0/10.1 hardware. Microsoft requires applications to query
D3D11_FEATURE_D3D10_X_HARDWARE_OPTIONS and specifically
ComputeShaders_Plus_RawAndStructuredBuffers_Via_Shader_4_x; support cannot be
assumed from the model name alone. See Microsoft's
downlevel compute documentation
and the
feature-query structure.
The current backend now:
- stores the device's actual feature level;
- queries the optional DirectCompute 4.x capability on FL10 hardware;
- compiles SM4.0 compute/vertex/pixel profiles on FL10 and SM5.0 on FL11;
- completely skips compute shaders and UAV particle buffers for fragment-only workloads, so those tests remain available even when the optional compute bit is absent;
- rejects FP64 SynthPeak and oversized dispatches cleanly on the downlevel path;
- records feature level, shader model, and DirectCompute availability with the saved driver metadata.
SM4 downlevel compute is constrained but real: only one compute UAV can be
bound, typed UAVs are unavailable, thread-group shared memory is limited to
16 KiB, and a group is limited to 768 threads. The original Stream shader uses
256 threads and one RWStructuredBuffer, so it fits this contract. N-body now
uses SV_GroupIndex, allowing its small configurations to compile for SM4;
SynthPeak FP32/INT32 also compiles, while FP16/FP64 remain excluded.
| Era | DirectX / FL | Shader Model | Relevant capability |
|---|---|---|---|
| Advanced raster | DX9.0c | SM 3.0 | VS/PS, dynamic branching; no compute/UAV |
| Unified raster | DX10 / FL10_0 | SM 4.0 | VS/PS/GS and optional DirectCompute 4.x through the D3D11 runtime |
| Refined unified | DX10.1 / FL10_1 | SM 4.1 | SM4.1 raster additions and optional DirectCompute 4.x |
| Full DX11 compute | DX11 / FL11_0 | SM 5.0 | Standard compute, richer UAV/TGSM/atomic/tessellation feature set |
| Explicit APIs | DX12 / Vulkan | SM 5.1+ / SPIR-V | Explicit command recording, queues, and modern resource control |
This does not mean every workload is safe to launch at its modern default on
a GT 120. First acceptance should use Stream/Particle Light or Medium and small
N-body/Render3D configurations. Large SynthPeak loops, GPU Burn/Stress, and
Volumetric need a separately versioned legacy_sm4 calibration to avoid Windows
TDRs. FP16, FP64, Vulkan, DX12, OpenGL 4.3, Fluid, and Cinematic Liquid are not
part of the GT 120 contract.
DX9 could be implemented as a separate raster backend, but it has no compute shader/UAV path and would require a new D3D9 device, presentation, timestamp, resource, and SM3 shader implementation. Its results could not share the Stream, N-body, or SynthPeak contracts. RenderDoc also does not support D3D9 capture, so it would break this project's fixed fifth-second capture workflow. Because the GT 120 already exposes FL10_0, D3D11 downlevel is both the more capable and the more comparable route.
The code and all SM4 profiles have been compiled on the development machine, but the optional DirectCompute bit, driver timing behavior, safe calibration, and fifth-second RenderDoc capture still require validation on the physical GT 120 before any result is called formal.
Three concepts are often conflated but serve distinct roles:
- Rendering pipeline — the GPU's processing stages (vertex → rasterisation → fragment → output, plus optional stages like geometry, tessellation, and compute). This defines what stages exist and how data flows between them.
- Shader Model (SM) — a hardware capability specification defining what code can run at each programmable stage. Higher SM versions unlock more instructions, longer programs, new memory access patterns, and new pipeline stages.
- Shading languages — the programming languages developers use to write shader code for each stage.
These three evolve together: a new pipeline stage (e.g., compute shader) requires new hardware capability (SM 5.0), which is then exposed through shading language features (HLSL [numthreads], GLSL layout(local_size_x)).
Each graphics API defines its own shading language:
| Language | API | Compilation | Target |
|---|---|---|---|
| HLSL (High-Level Shading Language) | DirectX 9–12 | fxc (SM 2.0–5.0) / dxc (SM 6.0+) |
DXBC / DXIL bytecode |
| GLSL (OpenGL Shading Language) | OpenGL / Vulkan | glslang / glslc |
OpenGL: driver compiles at runtime; Vulkan: pre-compiled to SPIR-V |
| MSL (Metal Shading Language) | Metal | Metal compiler | Apple IR |
| SPIR-V | Vulkan | Intermediate representation | Consumed by Vulkan driver, compiled to GPU-native ISA |
The compilation pipeline in this benchmark:
Vulkan: GLSL (.comp/.vert/.frag) ──→ glslc ──→ SPIR-V (.spv) ──→ Vulkan driver ──→ GPU ISA
DX12: HLSL (.hlsl) ──→ fxc ──→ DXBC bytecode ──→ DX12 driver ──→ GPU ISA
DX11: HLSL (.hlsl) ──→ fxc ──→ DXBC bytecode ──→ DX11 driver ──→ GPU ISA
OpenGL: GLSL (embedded strings) ──→ driver compiles at runtime ──→ GPU ISA
Despite using different languages, all backends compile down to the same GPU instruction set (ISA) for a given GPU. The shader logic is equivalent across all four backends — position update in the compute shader, point-sprite rendering in the vertex/fragment shaders. The only differences are syntax and API-specific boilerplate.
This is why cross-API comparisons in this benchmark are meaningful: the same algorithm runs through different API/driver paths to the same hardware, isolating the API and driver overhead from the shader workload itself.
The Shader Model version determines what a GPU can do, but different APIs expose this through different naming:
| SM | DX Feature Level | OpenGL Version | GLSL Version | Key Addition |
|---|---|---|---|---|
| SM 2.0 | 9_1 – 9_3 | 2.1 | 120 | Basic VS/PS, FP32 |
| SM 3.0 | — | 3.0 | 130 | Dynamic branching |
| SM 4.0 | 10_0 | 3.3 | 330 | Unified shaders, geometry shader, integer ops; optional DirectCompute 4.x through D3D11 |
| SM 4.1 | 10_1 | — | — | Gather4/MSAA additions; optional DirectCompute 4.x through D3D11 |
| SM 5.0 | 11_0 | 4.3 | 430 | Standard full compute/UAV feature set and tessellation |
| SM 5.1 | 11_1 / 12_0 | 4.5+ | 450 | Bindless-style resource indexing |
| SM 6.0+ | 12_0+ | — (Vulkan SPIR-V) | — | Wave intrinsics, ray tracing, mesh shaders |
The modern cross-API comparison matrix still starts at SM 5.0 / Feature Level 11_0 / OpenGL 4.3. A new, explicitly separated DX11-downlevel contract can include FL10_0/SM4 hardware when its driver exposes DirectCompute 4.x. Those results must record the feature level and shader profile and must not be mixed with Vulkan/DX12/OpenGL 4.3 coverage claims.
These results quantitatively demonstrate that modern APIs (Vulkan, DX12) are essential for realising the full compute potential of AMD hardware. The OpenGL compute path carries significant driver overhead on AMD, particularly on legacy architectures, reinforcing the industry trend towards explicit, low-overhead graphics APIs. Choosing the right API is more impactful than optimising application code when driver-level overhead dominates.
The RX 9070 XT (RDNA 4) is the most extreme example: OpenGL compute takes 2.612 ms vs Vulkan's 0.033 ms — a 79× penalty — the largest ratio observed in any GPU tested. Despite being AMD's newest architecture, the OpenGL compute dispatch overhead has not improved from RDNA 2 levels (~2.7 ms), confirming that AMD's driver team has deprioritised OpenGL compute optimisation.
Additionally, API performance rankings are architecture-dependent: DX11 leads on modern NVIDIA GPUs but falls behind Vulkan on older Maxwell hardware due to compute–render synchronisation costs. On the RX 9070 XT, DX11 and Vulkan are nearly tied (1,774 vs 1,751 FPS), reflecting AMD's less mature DX11 optimisation path. This underscores that no single API is universally optimal — the best choice depends on the target hardware generation and driver maturity.
The Mac Pro (Late 2013) contains two identical AMD FirePro D700 GPUs (GCN 1.0, Tahiti XT). Two concepts must be kept separate on this machine: selecting one adapter independently, and making both adapters collaborate on one benchmark. Earlier report revisions conflated them.
Before the D700 measurements, four terms are easy to mix up. AFR and SFR are how two GPUs share work inside one application. CrossFire (AMD) and SLI (NVIDIA) are vendor multi-GPU brands / driver frameworks that historically packaged those modes for games—usually as an opaque driver profile rather than an app-controlled API contract.
| Term | What it means | How work is split | Typical cost / risk |
|---|---|---|---|
| AFR (Alternate Frame Rendering) | GPUs take turns rendering whole frames (GPU0 → frame N, GPU1 → frame N+1, …) | Frame-level parallelism; both GPUs can overlap if the queue is deep enough | Extra frames-in-flight; higher input latency; bad for strongly stateful frame-to-frame dependencies unless state is copied |
| SFR (Split Frame Rendering) | Both GPUs render parts of the same frame (commonly left/right halves, or other screen partitions), then one present path composites | Intra-frame data parallelism; one logical image | Cross-GPU copy/composite every frame; can lose to AFR when the transfer cost dominates light shaders |
| CrossFire | AMD’s multi-GPU product/driver branding (later “mGPU” naming varied by generation) | Historically often AFR via driver profiles; some eras also SFR / hybrid modes | App may get no control and no honest utilization signal; profile-dependent and largely retired for modern DX12/Vulkan games |
| SLI | NVIDIA’s multi-GPU product/driver branding | Same idea family as CrossFire—mostly driver-managed AFR for DX11-era titles | Same opacity problem; modern NVIDIA dual-GPU gaming support is effectively discontinued |
Key distinctions for this benchmark:
- Mode vs brand. AFR/SFR describe the scheduling. CrossFire/SLI name the vendor stack that might implement a mode for you. Saying “CrossFire is on” does not prove the app is doing explicit AFR or SFR—only that the driver claims multi-GPU participation.
- Explicit engine path vs implicit driver path. Mangekyo’s validated collaboration is application-controlled DX12 linked-adapter work (
NodeMask, per-node queues/resources,--multi-gpu afr|sfr). That is not the same as hoping DX11 CrossFire/SLI profiles accelerate a custom engine. On this D700 FireGL stack, AMD AGS still saw two physical cards but returnedcrossfireAPI=0/ one active GPU for both explicit and driver AFR requests—so classic CrossFire was unavailable even when two adapters existed. - Why Plasma can use AFR, while Particle prefers SFR. Plasma/GPU Burn has no persistent simulation state between frames, so alternating whole frames is meaningful and measured ~2× on DX12 AFR (with ≥4 frames in flight, RenderDoc off). Original Particle’s frame N+1 depends on frame N particle state; AFR would require a full cross-node state copy every frame or would change the test semantics. The product direction for Particle is therefore fixed-total-count SFR/data-parallel split of one frame, not AFR.
- SLI is not “validated” here. The DX12 linked-node code is vendor-neutral in the sense Microsoft’s linked-GPU samples describe homogeneous CrossFire/SLI topologies, but this report only accepts measured D700 DX12 AFR/SFR results. No NVIDIA SLI system has been validated; the GUI must not claim SLI support.
In short: AFR = alternate whole frames; SFR = split one frame; CrossFire/SLI = vendor multi-GPU brands that may hide either mode behind the driver. This project only scores modes it controls and measures.
| API | Independent adapter selection | Explicit two-GPU collaboration | D700 result |
|---|---|---|---|
| Vulkan | Two VkPhysicalDevice objects; GPU #2 can be selected directly |
Physical-device-group masks | Functional, but the only graphics-capable family exposes one VkQueue, so tested AFR submissions serialize and do not scale |
| DirectX 12 | Driver-dependent: 25.20.14020.10001 exposed two distinct DXGI LUIDs but routed both to the primary GPU; 27.20.14540.15002 exposes one linked adapter |
Linked-adapter device with two explicit nodes and NodeMask control |
Working AFR (~2×) and workload-dependent SFR (1.47× at 128 steps) |
| DirectX 11 | The old driver accepted two DXGI LUIDs but routed both to the primary GPU; the current topology exposes one logical adapter and AGS still sees both physical D700s | Standard D3D11 has no node control; AGS can request explicit or driver-managed CrossFire when the driver exposes it | Windowed implicit AFR did not scale; AGS returned crossfireAPI=0 and one active GPU for both AFR modes |
| OpenGL | The AMD WGL extension enumerates both D700 GPU IDs | WGL_AMD_gpu_association would permit per-GPU off-screen contexts and cross-context framebuffer blit |
Driver exposes the interface but refuses to create an associated context for GPU #2; explicit SFR is unavailable |
DXGI LUID matching remains necessary because factory instances can reorder adapters, and VendorId + DeviceId + SubSysId is not a unique identity. The archived March run on 25.20.14020.10001 recorded two distinct D700 LUIDs and successfully created DX11/DX12 devices through both handles. Task Manager nevertheless showed both D3D runs executing on the primary physical GPU, so those are valid API runs but not independent second-card measurements. After reinstalling 27.20.14540.15002, the current probe sees one D700 DXGI LUID; ID3D12Device::GetNodeCount() exposes its two physical nodes. D3D11 cannot address those nodes independently, but D3D12 can: the benchmark now expands a linked adapter into selectable node rows and binds all ordinary single-GPU objects and resources to the selected node mask.
A July 21 run-all check on the current driver exposed a probe-merging and routing bug: Vulkan correctly returned two VkPhysicalDevice objects, but the synthetic second row copied the first row's DXGI identity while the ordinary DX12 backend left NodeMask=0. Consequently both labelled DX12 runs executed on linked-adapter node 0; the two 65K Particle results (82.610 and 82.723 GB/s) were not independent card measurements. The corrected probe derives two rows from GetNodeCount() and correlates the two Vulkan devices with those rows. The verified capability table is now D700 #1 = Vulkan/DX12/DX11/OpenGL and D700 #2 = Vulkan/DX12. For ordinary DX12, the command queue, ResizeBuffers1 placement, command lists, root signatures, PSOs, descriptor/query heaps, timestamps, and particle/fluid resources all use the row's explicit mask (0x1 or 0x2). A temporary 4M Particle validation produced 122.8 FPS / 78.66 GB/s on node 0 and 120.1 FPS / 80.14 GB/s on node 1. Each result persists dx12Node, dx12NodeMask, and dx12LinkedNodes in workloadConfig. This correction does not retroactively assert that the old driver's two LUIDs were fabricated; it records that the old physical routing and the new linked-node topology are different cases.
The Windows GUI now exposes Multi-GPU: Off / AFR / SFR with an adjacent information icon. It is enabled only for a Custom run whose sole API is DX12 and whose workload is Plasma; selecting AFR or SFR forwards the matching --multi-gpu argument and disables RenderDoc, while changing to an unsupported combination resets the mode to Off. The code is not AMD-specific: a NVIDIA SLI configuration could use the same path only when its driver presents the pair as one DX12 linked adapter with GetNodeCount() >= 2 and the required cross-node sharing support. Microsoft's official D3D12 linked-GPU sample explicitly describes the linked homogeneous case as CrossFire/SLI, but this benchmark has not yet validated an NVIDIA system and therefore does not claim SLI support.
History now renders the saved execution contract explicitly in a Mode column: Single, Headless, AFR ×2, or SFR ×2. Legacy DX11/Vulkan AFR experiments are labelled unverified/experimental from their persisted control tags. The headless boolean was already serialized, but the generic comparison grouping previously omitted it and could place windowed and headless stream_v1 rows together. Comparison groups now include execution mode, and pairwise score deltas require matching headless state.
The final DX11 probe used the official AMD AGS 6.3.1 library rather than inferring CrossFire activity from Task Manager. agsInitialize enumerated both FirePro D700s. agsDriverExtensionsDX11_CreateDevice also succeeded in both AGS_CROSSFIRE_MODE_EXPLICIT_AFR and AGS_CROSSFIRE_MODE_DRIVER_AFR, but both returned extensionsSupported.crossfireAPI=0 and crossfireGPUCount=1. Thus this FireGL driver exposes two physical adapters but does not activate or expose DX11 CrossFire for the application; there is no useful AGS integration to retain in the benchmark. AMD defines explicit AFR as the no-profile path and crossfireGPUCount as the number of GPUs active for the app in the official AGS DX11 API.
The OpenGL SFR probe used the only relevant explicit Windows mechanism, WGL_AMD_gpu_association. FireGL 20.45.40.15 reports two IDs and maps the visible context to ID 1, but ID 2 returns no renderer string. Both wglCreateAssociatedContextAttribsAMD for a 4.3 core context and legacy wglCreateAssociatedContextAMD return NULL, including a retry with the visible context unbound; the ICD also leaves GetLastError at zero. Consequently the second D700 cannot receive OpenGL commands. The planned left/right rendering plus wglBlitContextFramebufferAMD composition cannot begin, so the prototype and CLI enablement were reverted rather than falling back to single-GPU rendering. The intended mechanism and its requirement for a valid GPU-associated context are defined by the Khronos WGL_AMD_gpu_association specification.
The first correct AFR vertical slice is limited to the windowed Plasma/GPU Burn workload (--multi-gpu afr). It alternates complete frames between DX12 node masks 0x1 and 0x2 and uses a separate queue and frame resources for each node.
| Backend / configuration | 16 shader steps | 128 shader steps | Interpretation |
|---|---|---|---|
| DX12 single GPU | ~113 FPS | ~19 FPS | Baseline |
| DX12 AFR, 2 frames in flight | ~109 FPS | — | Too little queue depth; each node has only one reusable slot |
| DX12 AFR, 4 frames in flight | 222–230 FPS | 39–40 FPS | 1.96–2.05×; both nodes overlap |
| Vulkan single / device-group AFR | ~102 / ~104 FPS | ~17 / ~17 FPS | No material gain; shared graphics VkQueue serializes work |
| DX11 single / implicit-driver AFR request | ~120 / ~120 FPS | ~20 / ~20 FPS | No gain; driver-managed physical split unverified |
The program therefore forces at least four frames in flight for AFR. Utilisation analysis is reported as GPU-equivalent work across the two devices: the validated DX12 run reached approximately 200% aggregate (about 100% per GPU), whereas Vulkan and DX11 remained near 100% aggregate.
Additional Vulkan isolation tests moved vkQueuePresentKHR from queue family 0 to the independently present-capable transfer family 2. This removed presentation operations from the sole graphics queue without changing the result: 16-step AFR remained about 103 FPS, while the 128-step single/AFR pair remained 17/17 FPS with approximately 100% aggregate GPU-equivalent utilisation. A paired-acquire experiment intended to enqueue both device-mask submissions before either present instead blocked indefinitely in AMD 20.45.40.15's second vkAcquireNextImage2KHR call on the LOCAL-only swapchain and was reverted.
Both D700 multi-instance device-local heaps expose peer access as VK_PEER_MEMORY_FEATURE_COPY_DST_BIT only in either direction—no peer copy-source or generic reads. A secondary GPU can therefore write a destination allocated for the primary GPU, but the primary cannot directly pull from secondary-local output. The remaining single-VkDevice experiment used the compute-only family's two queues and an R8G8B8A8_UNORM storage-capable swapchain: alternate device masks dispatched both Plasma layers directly into each GPU's LOCAL swapchain image. Runtime capability checks, compilation and a RenderDoc-disabled three-second run all succeeded, including clean workload completion. The result was nevertheless only about 119 FPS / 8.18 ms at 16 steps and approximately 98% aggregate GPU-equivalent utilisation (49% average per GPU), proving that the driver still did not overlap alternate frames. Because the compute shader also changes the pipeline contract relative to fragment Plasma, this small throughput difference is not evidence of AFR scaling. The implementation, shader variant and separate score identity were reverted in full. Only a future, separately scoped two-VkDevice external-memory/external-semaphore design remains unexplored. The peer-memory flag meanings are defined by Khronos.
RenderDoc is not disabled globally. Single-GPU runs still use the normal 15-second run and fifth-second capture. It is automatically disabled for multi-GPU AFR and SFR because merely injecting RenderDoc on this old D700 stack produced false ~2000 FPS output, missing GPU timestamps and a long queue drain at process exit. SFR has not independently proved capture-safe. Multi-GPU scoring must use native timing plus GPUView/PIX-style diagnostics; a one-node reproduction may still be captured separately for shader debugging.
DX12 can alternate particle frames at the API level, but Original Particle is stateful: frame N+1 consumes the positions and velocities produced by frame N. Copying that full state across nodes every frame would add synchronization and transfer cost, while keeping one independent state per node changes the simulation and is not a valid acceleration. DX11 offers only implicit driver AFR and cannot solve this dependency explicitly. Consequently, particle AFR is not a product path.
The meaningful two-GPU particle design is SFR/data parallelism: keep one fixed total particle count, split the particle range between both GPUs, simulate and draw both partitions for the same frame, then composite once. That measures one accelerated test rather than two tests run side by side. Plasma is the lower-risk SFR prototype because it has no persistent simulation state: render the left and right halves into node-local targets, copy one half to the presenting node, then compose and present. Complexity is moderate rather than trivial—cross-node resource visibility, two fences, copy/composition and failure fallback must all be handled.
The version 0.2.2 runtime probe reports D3D12_CROSS_NODE_SHARING_TIER_1 on the D700 linked adapter. This is sufficient for explicit cross-node CopyBufferRegion, CopyTextureRegion and CopyResource operations when the shared resource is the copy destination, so a Plasma half-frame copy/composite prototype is technically viable. It does not promise free peer bandwidth or cross-node render-target use; the copy cost must be measured before SFR is accepted. See Microsoft's D3D12_CROSS_NODE_SHARING_TIER contract.
The prototype is now implemented as --multi-gpu sfr / --sfr. Both nodes use the same frame time and shader contract. Node 0 shades the left half directly into the swap-chain buffer; node 1 shades the right half into a node-local full-size render target so SV_Position and UV values remain identical to the single-GPU image. Node 1 then copies only the 640x720 right region into a 1,843,200-byte node-0-owned shared buffer. A cross-node fence lets node 0 consume that buffer, copy it into the right side of the back buffer, and perform the only Present. This is one cooperative frame, not two benchmark instances.
The old D700 driver rejects a cross-adapter committed resource with E_INVALIDARG; using the explicit shared-heap plus placed-resource pattern from Microsoft's heterogeneous multi-adapter sample succeeds. Two and four frames in flight produce the same result, ruling out a shallow frame ring. Short isolated comparisons on driver 27.20.14540.15002 were:
| Plasma fixed steps | Single DX12 | DX12 SFR | Scaling | Interpretation |
|---|---|---|---|---|
| 16 | 112 FPS / 8.17 ms | 69 FPS / 14.33 ms | 0.62x | Tier-1 transfer/composition dominates |
| 128 | 19 FPS / 49.75 ms | 28 FPS / 35.11 ms | 1.47x | half-frame shading repays the fixed copy cost |
SFR therefore works and accelerates sufficiently heavy Plasma frames, but it is not a universal optimisation and remains slower than the measured 128-step AFR result of about 39 FPS. Results are isolated under ..._sfr2; RenderDoc remains disabled for SFR until linked-node capture is independently trustworthy. Microsoft's Tier-1 contract permits only cross-node copies with the shared resource as destination, which is why the secondary GPU cannot render directly into the presenting resource. See the official cross-node sharing tiers and multi-adapter copy example.
| Observation | Explanation |
|---|---|
| DX11 > DX12 > Vulkan > OpenGL in FPS (hardware GPU, simple workload) | Per-frame CPU overhead: DX11 (0.008 ms) < DX12 (0.088 ms) < Vulkan (0.183 ms) < OpenGL (0.322 ms) |
| All four APIs have similar GPU execution times (0.065–0.104 ms) | The GPU-side workload is identical; only CPU-side driver overhead differs |
| OpenGL has fastest GPU time but lowest FPS | WGL swap path and state-machine validation overhead dominate CPU time |
| DX12 > DX11 in FPS (WARP software renderer) | DX11's implicit driver layer is pure overhead when there is no GPU hardware to optimise for |
| DX12 has lowest GPU time on hardware | Slightly more efficient GPU command scheduling, but CPU overhead negates the advantage at low complexity |
| dGPU ~37× faster than iGPU | Bandwidth-bound workload; ratio matches memory bandwidth difference |
| Device-local 35× faster than host-visible on dGPU | PCIe round-trips dominate compute shader memory access patterns |
| Host-visible = Device-local on iGPU | Unified memory architecture — no PCIe hop |
| WARP ~80× slower than hardware GPU | Expected for CPU-based software rasterisation |
| WARP can appear as a Vulkan device via Dozen | Mesa Dozen (Vulkan→D3D12) wraps WARP; only visible when no hardware Vulkan ICD is present |
| OpenGL cannot select GPU on Windows | No standard API; OS assigns GPU. Linux provides DRI_PRIME as a workaround |
| DX11 timestamps fail where Vulkan/DX12 succeed | DX11's Disjoint flag discards data on GPU clock changes; Vulkan/DX12 have no such mechanism |
| RDNA 2 iGPU (0.56 TFLOPS) beats HD 5770 (1.36 TFLOPS) by 42–71% | VLIW5 utilisation losses, immature compute scheduler, and stale drivers reduce TeraScale 2's effective throughput well below its theoretical peak |
| TFLOPS is a poor predictor of compute shader performance | Architectural efficiency (SIMD vs VLIW), driver maturity, and compute scheduler design dominate raw ALU count |
| HD 5770 would likely win a gaming benchmark | Traditional rasterisation relies on fixed-function units (TMUs, ROPs) where the HD 5770 has 4–6× more hardware than the iGPU |
| Fast GPUs report inflated "render time" in windowed mode | Vulkan's COLOR_ATTACHMENT_OUTPUT_BIT timestamp includes swapchain semaphore wait; on fast GPUs (9070 XT), ~90% of reported render time is idle wait |
| Increasing swapchain BufferCount does not fix timestamp pollution | The semaphore wait is inherent to the Vulkan presentation model, independent of buffer pool size |
| Headless mode achieves 10× FPS over windowed mode | Removing swapchain/render/present eliminates presentation overhead; all APIs converge to ~0.034 ms compute time |
| DX11 headless requires workarounds for timestamp queries | Without Present() as a frame boundary, DX11's timestamp pipeline produces garbage values; sanity filtering discards ~3–4% of samples |
OpenGL headless requires periodic glFinish() on AMD |
AMD's driver doesn't actively process commands for hidden windows; periodic full sync every 16 frames restores timestamp availability |
| 3DMark Unlimited ≠ headless compute | 3DMark Unlimited renders offscreen (full pipeline); this benchmark's headless mode skips rendering entirely (compute only) |
| RX 9070 XT (RDNA 4) has 4.1× better per-CU compute than RX 6600 XT (RDNA 2) | 2× the CU count (64 vs 32), 1.15× higher clocks; the remaining ~1.8× is architectural improvement in scheduler, cache, and driver codegen |
| 9070 XT outperforms 80-CU RX 6900 XT in headless compute (21K vs 16K FPS) | RDNA 4's per-CU efficiency (~1.7× higher) overcomes the 1.25× CU disadvantage; on a per-dollar basis, 9070 XT delivers 1.9× the FPS/$ of RTX 5090 |
| 16M particles reveals primitive throughput bottleneck | RX 6900 XT (218 FPS) beats 9070 XT (111 FPS) at 16M — 80 CUs provide 1.25× more rasterisation hardware; RTX 5090 (612 FPS) leads due to 170 SMs |
| 9070 XT compute scales 59× for 16× particles (1M → 16M) | Super-linear scaling due to GPU underutilisation at 1M; at 16M the GPU is fully saturated |
| 9070 XT headless achieves 21K FPS across all APIs | All APIs converge to 0.034 ms compute; windowed mode's 1.7K FPS is 12× slower due to presentation overhead |
These results demonstrate that API overhead, memory placement, and hardware architecture all significantly affect GPU compute performance — and that the optimal configuration depends on workload complexity and hardware topology.
The Adreno 640 tested in this benchmark shares a direct lineage with the Radeon GPUs it is compared against. This section traces the corporate and technical connections from ATI Technologies through AMD to Qualcomm's Adreno mobile GPU division.
| Year | Event |
|---|---|
| 1985 | Array Technology Inc. (ATI) founded in Markham, Ontario, Canada |
| 1987 | First product: EGA Wonder — ISA graphics card for IBM PCs |
| 1991 | ATI enters the dedicated 2D accelerator market (Mach 8, Mach 32) |
| 1996 | 3D Rage — ATI's first 3D-capable GPU. Competed with 3dfx Voodoo and S3 ViRGE |
| 2000 | Radeon DDR (R100) — ATI's first GPU under the Radeon brand. Hardware T&L, competed with NVIDIA GeForce 2 |
| 2002 | Radeon 9700 Pro (R300) — first DirectX 9 GPU, SM 2.0, outperformed GeForce FX. Widely considered ATI's finest moment |
| 2006 | AMD acquires ATI Technologies for $5.4 billion. ATI's GPU division becomes AMD Graphics |
| 2008 | AMD sells its mobile GPU division (Imageon) to Qualcomm for $65 million |
| 2009 | Qualcomm renames Imageon to Adreno (anagram of "Radeon"). First product: Adreno 200 in Snapdragon QSD8250 |
| 2013 | AMD rebrands consumer GPUs from "Radeon HD" to "Radeon R" series, then later "Radeon RX" |
| 2017 | AMD launches Vega architecture (GCN 5). "ATI" name fully phased out from all products |
| 2020 | AMD launches RDNA 2 (RX 6000 series). Qualcomm launches Adreno 660 (Snapdragon 888) |
| 2024 | Qualcomm launches Snapdragon X Elite with Adreno X1 GPU for Windows on ARM laptops |
| 2025 | AMD launches RDNA 4 (RX 9070 XT). Qualcomm's Adreno GPUs power the majority of Android devices and Windows on ARM PCs |
ATI's Imageon was a low-power mobile GPU line designed for handheld devices and embedded systems (PDAs, early smartphones). When AMD acquired ATI in 2006, Imageon became part of AMD's portfolio but was considered non-core — AMD's focus was on discrete desktop/laptop GPUs (Radeon) and professional workstation GPUs (FirePro).
In 2008, AMD divested the Imageon mobile GPU division to Qualcomm for $65 million — a fraction of the $5.4 billion AMD paid for all of ATI. Qualcomm integrated Imageon into its Snapdragon SoC platform and renamed it Adreno — an anagram of "Radeon" that preserves the ATI heritage while establishing a distinct brand.
Adreno's architecture has since diverged significantly from Radeon. By the Adreno 600 series (2018), the GPU shares no meaningful silicon design with contemporary Radeon GPUs — the instruction set, memory hierarchy, shader core layout, and driver stack are entirely Qualcomm-designed. However, foundational concepts from ATI's Imageon era (TBR, tile-based rendering optimisations, unified shader architecture for mobile power budgets) persist in Adreno's design philosophy.
TBR vs IMR — and the modern reality. The rendering architecture inherited from Imageon is Tile-Based Rendering (TBR): the screen is divided into small tiles (e.g. 16×16 or 32×32 pixels), and each tile is rendered entirely within fast on-chip tile memory before being written to DRAM — minimising bandwidth-hungry framebuffer read/write traffic. Traditional desktop GPUs use Immediate Mode Rendering (IMR): triangles are rasterised and written to the framebuffer in submission order, relying on high memory bandwidth rather than on-chip buffering. In practice, however, no modern GPU is purely TBR or purely IMR. Mobile GPUs like Adreno and Apple's M-series use TBDR (Tile-Based Deferred Rendering), which adds deferred visibility testing (Hidden Surface Removal) per tile to eliminate overdraw before shading. Meanwhile, desktop GPUs have adopted tile-like techniques: NVIDIA's Maxwell and later use a tiling rasteriser that groups fragments into screen-space tiles for more efficient L2 cache usage; AMD's RDNA series introduced binning passes (a form of tiling) to reduce bandwidth — the Infinity Cache discussed throughout this report is particularly effective when combined with RDNA's binning, as tiles that fit in cache avoid VRAM round-trips entirely. The distinction today is a spectrum: mobile GPUs lean tile-heavy with optional immediate fallback for complex geometry, while desktop GPUs are fundamentally immediate-mode but borrow tiling techniques for bandwidth efficiency.
ATI Technologies (1985)
├── Radeon R100 (2000)
│ └── Radeon 9700 (R300, 2002) — first DX9 GPU
│ └── Radeon X1800 (R520, 2005) — last ATI-only design
│
├── Imageon (mobile GPU line, 2002–2008)
│ └── [Sold to Qualcomm, 2008]
│ └── Adreno 200 (2009) — renamed from Imageon
│ └── Adreno 3xx/4xx/5xx/6xx
│ └── Adreno 640 (2019) ← TESTED: Xiaomi Pad 5 / Snapdragon 860
│ └── Adreno X1 (2024) — Snapdragon X Elite for WoA
│
└── [AMD acquires ATI, 2006]
└── AMD Radeon
├── HD 5770 (TeraScale 2, 2009) ← TESTED
├── FirePro D700 (GCN 1.0, 2013) ← TESTED
├── RX 580 (GCN 4, 2017) ← TESTED
├── Vega FE (GCN 5, 2017) ← TESTED
├── RX 6600 XT / 6900 XT (RDNA 2, 2020–2021) ← TESTED
└── RX 9070 XT (RDNA 4, 2025) ← TESTED
The Adreno 640 in this benchmark is, in a historical sense, a distant cousin of the Radeon GPUs it is compared against. Both trace their origins to ATI Technologies, but their architectures diverged completely after the 2008 sale to Qualcomm. Comparing them side-by-side in the same benchmark highlights:
-
How far mobile GPUs have come: The Adreno 640 (a 2019 mobile GPU in a tablet SoC) can run the same Vulkan 1.1 compute + render pipeline as desktop GPUs, achieving 106 FPS — slower than a discrete desktop GPU, but functional and measurable with identical code.
-
The power/performance trade-off: The Adreno 640 operates within a ~3W thermal envelope (tablet SoC), while the RX 9070 XT consumes ~300W. The 9070 XT is ~17× faster (1,751 vs 106 FPS), but uses ~100× more power — making the Adreno 640 significantly more performance-per-watt efficient for this workload.
-
Driver maturity gap: Qualcomm's Windows Vulkan driver is relatively young (first WoA devices shipped 2023), while AMD's Radeon Vulkan drivers have been refined since 2016. This is reflected in the Adreno 640's higher render times and less efficient presentation path.
First-ever inclusion of a mobile Qualcomm GPU in this benchmark suite. The Adreno 640 runs natively on Windows 11 ARM64 via Qualcomm's WoA Vulkan/DX12/DX11 drivers — no emulation or translation layer for the GPU workload itself.
The test device is a Xiaomi Pad 5 — an Android tablet originally shipping with MIUI 12.5 based on Android 11. Through community-developed custom firmware (Project Renegade / Windows on ARM for Snapdragon 855/860 tablets), Windows 11 ARM64 was installed on this device, replacing the Android operating system entirely. The tablet boots directly into Windows 11, with Qualcomm-provided WDDM drivers exposing the Adreno 640 as a standard Windows GPU — complete with Vulkan, DirectX 12, and DirectX 11 support.
The motivation for testing this device is rooted in the ATI → AMD → Qualcomm lineage documented in the previous section. The Adreno GPU traces its origin to ATI's Imageon mobile GPU division, which AMD sold to Qualcomm in 2008. Qualcomm renamed Imageon to Adreno — an anagram of "Radeon" — and has since developed it into a fully independent architecture. While Adreno shares no silicon or ISA with modern Radeon GPUs, the historical connection means this benchmark suite now covers both branches of ATI's GPU family tree: the Radeon line (TeraScale 2 → GCN → RDNA 4) and the Adreno line that diverged in 2008.
Including the Adreno 640 also provides unique data points not available from any desktop GPU:
- Mobile vs desktop GPU scaling: How does a 2-5W tablet SoC GPU compare against 75–300W discrete desktop GPUs running the exact same compute + render workload?
- Qualcomm driver maturity: Qualcomm's Windows GPU drivers are relatively new (first WoA devices shipped 2023). How do they perform compared to AMD and NVIDIA's decade-old Windows driver stacks?
- ARM64 vs x64 binary translation: The same benchmark compiled as ARM64 (native) and x64 (Prism emulation) on identical hardware reveals the exact CPU-side overhead of Microsoft's binary translation layer — with GPU times serving as a constant control variable.
| Component | Specification |
|---|---|
| Device | Xiaomi Pad 5 (nabu) — Android tablet running Windows 11 ARM64 via Renegade Project / Port-Windows-11-Xiaomi-Pad-5 |
| Previous OS | Android 15 (HyperOS 2) |
| Current OS | Windows 11 Pro ARM64 (NT 10.0.26100) |
| SoC | Qualcomm Snapdragon 860 (SM8150-AC) — a binned Snapdragon 855+ |
| CPU | Kryo 585: 1× A77 @ 2.96 GHz (prime) + 3× A77 @ 2.42 GHz (performance) + 4× A55 @ 1.80 GHz (efficiency) |
| GPU | Qualcomm Adreno 640, ~585 MHz boost, 384 ALUs |
| RAM | 12 GB LPDDR4X (shared between CPU and GPU) |
| VRAM | Shared (reported as 1 MB by driver — a Qualcomm WDDM driver reporting limitation, not actual VRAM size) |
| Vulkan | 1.1.276 (Qualcomm proprietary driver, build 2023-10-23) |
| DX12 | Feature Level 12_1 (driver 27.20.2060.0) |
| DX11 | Supported (driver 27.20.2060.0) |
| OpenGL | Not supported — Qualcomm's Windows driver does not expose desktop OpenGL. Microsoft's OpenCL/OpenGL Compatibility Pack provides only GL 3.3 via Mesa-on-DX12 translation, below this benchmark's GL 4.3 requirement. (The same hardware supports OpenGL ES 3.2 on Android, and the open-source Mesa Freedreno driver achieves full OpenGL 4.6 on Linux.) |
| Display | 11" 2560×1600 IPS 120 Hz (benchmark runs at 1280×720) |
| Thermal | Passive cooling only (tablet form factor, no fan) |
| Resolution | 1280 × 720 |
| V-Sync | OFF |
| # | API | Avg FPS | Compute (ms) | Render (ms) | Total GPU (ms) | Frame Time (ms) | GPU Util | Bottleneck |
|---|---|---|---|---|---|---|---|---|
| 1 | DX12 | 116.7 | 3.281 | 4.767 | 8.053 | 8.57 | 94% | GPU-bound |
| 2 | Vulkan | 105.8 | 3.299 | 5.704 | 9.002 | 9.45 | 95% | GPU-bound |
| 3 | DX11 | 83.2 | 4.374 | 4.579 | 11.775 | 12.02 | 98% | GPU-bound |
1. DX12 is the fastest API on Adreno 640
Unlike desktop GPUs where DX11 or Vulkan often lead, DX12 is the clear winner on this mobile GPU. This suggests Qualcomm's DX12 driver path is more optimised than their Vulkan or DX11 paths — plausible given that DX12 is the primary API for Windows on ARM gaming and application compatibility.
2. Vulkan render time is inflated
Vulkan's render time (5.704 ms) is 20% higher than DX12's (4.767 ms), despite both rendering the same workload. This likely reflects immaturity in Qualcomm's Vulkan presentation/swapchain path on Windows, similar to the semaphore wait pollution observed on desktop GPUs (Section 6) but more pronounced.
3. DX11 has the highest compute overhead
DX11 compute time (4.374 ms) is 33% higher than Vulkan/DX12 (~3.3 ms). Combined with a large total GPU time (11.775 ms), DX11 is clearly the least efficient path. The implicit driver overhead that helps DX11 on mature desktop drivers (NVIDIA, AMD) does not translate to Qualcomm's younger driver stack.
4. All APIs are GPU-bound at 95%+ utilisation
The Adreno 640 is fully saturated at 1M particles — there is no CPU overhead headroom. This contrasts sharply with desktop GPUs where most APIs are CPU-bound at this particle count (e.g., RTX 5090 at 33% GPU utilisation). The mobile GPU's lower compute throughput means it hits the GPU-bound regime much earlier.
| GPU | Best API | Best FPS | Compute (ms) | Total GPU (ms) | vs Adreno 640 |
|---|---|---|---|---|---|
| RTX 5090 | DX12 | 5,603 | 0.014 | 0.065 | 48× faster |
| RX 9070 XT | DX11 | 1,774 | 0.047 | 0.542 | 15× faster |
| RX 6600 XT | DX12 | 1,834 | 0.190 | 0.406 | 16× faster |
| Vega FE | DX12 | 1,716 | 0.219 | 0.452 | 15× faster |
| RX 580 | DX12 | 912 | 0.362 | 0.930 | 8× faster |
| GTX 970 | Vulkan | 719 | 0.434 | 1.098 | 6× faster |
| FirePro D700 | Vulkan | 555 | 0.589 | 1.473 | 5× faster |
| Radeon iGPU (2 CU) | DX12 | 324 | 1.480 | 2.953 | 3× faster |
| HD 5770 | OpenGL | 188 | 1.794 | 4.818 | 1.6× faster |
| Adreno 640 | DX12 | 117 | 3.281 | 8.053 | 1.0× (baseline) |
| WARP on 9800X3D¹ | DX12 | 86.6 | 1.034 | 11.059 | 0.7× (slower) |
| WARP on 7600¹ | DX12 | 62.2 | 1.810 | 15.181 | 0.5× (slower) |
| WARP on SD 860 (native)¹ | DX12 | 4.2 | 30.416 | 187.321 | 0.036× (slower) |
| WARP on SD 860 (x64 emulated)¹ | DX12 | 2.9 | 56.804 | 338.073 | 0.025× (slower) |
¹ WARP is a CPU software renderer — performance depends entirely on the host CPU. The Ryzen 7 9800X3D (8-core, 96 MB 3D V-Cache) is 39% faster than the Ryzen 5 7600 (6-core, 32 MB L3), and both x86 desktop CPUs are dramatically faster than the Snapdragon 860 (1+3+4 ARM cores). On the SD 860, WARP was tested twice: the native ARM64 build (enumerated as "Microsoft Basic Render Driver") achieves 4.2 FPS, while the x64 emulated build (enumerated as "Microsoft WARP (CPU Software Renderer)", running through Prism translation) achieves only 2.9 FPS — a 45% penalty from binary translation on an already slow CPU. Even the native WARP result is 28× slower than the Adreno 640 hardware GPU on the same device, demonstrating the importance of hardware-accelerated GPU compute on mobile SoCs.
The Adreno 640 sits between the HD 5770 (a 2009 discrete desktop GPU) and the WARP software renderer in absolute performance. It outperforms the fastest WARP result (9800X3D) by 35%, confirming it is a real hardware GPU despite its mobile origins.
| API | Adreno 640 (WoA) | Desktop GPU (x64) | Notes |
|---|---|---|---|
| Vulkan | 1.1 (native driver) | 1.3+ | Qualcomm provides a native ARM64 Vulkan ICD |
| DX12 | FL 12_1 (native driver) | FL 12_1–12_2 | Full native support via WDDM driver |
| DX11 | Supported (native driver) | Supported | Full native support |
| OpenGL | 3.3 max (compatibility pack) | 4.6 | No native desktop OpenGL; Microsoft's compatibility pack (Mesa → DX12 translation) maxes out at GL 3.3, below the 4.3 required by this benchmark |
| Metal | N/A | N/A (macOS only) | — |
The lack of OpenGL 4.3 support means the Adreno 640 cannot run this benchmark's OpenGL backend. This is a platform limitation, not a hardware one — the same Adreno 640 on Android supports OpenGL ES 3.2 (roughly equivalent to desktop GL 4.3 in capability), and on Linux the open-source Mesa Freedreno driver achieves full OpenGL 4.6 on Adreno 600-series hardware.
Windows on ARM runs native ARM64 binaries at full speed, but can also execute x86/x64 applications through Microsoft's built-in binary translation layer (Prism). This section compares the same benchmark compiled as ARM64 vs x64 on identical hardware.
Windows 11 on ARM includes Prism, a binary translation layer that converts x86/x64 instructions to ARM64 at runtime. This allows unmodified x64 Windows applications to run on ARM hardware with a performance penalty. Prism translates code JIT (just-in-time), caching translated blocks for reuse.
Key characteristics:
- CPU code is translated: all C++ application logic (particle initialisation, frame loop, API calls) runs through the translation layer
- GPU code is NOT translated: shaders (SPIR-V, HLSL, GLSL) execute natively on the GPU regardless of the host binary's architecture
- API calls pass through: Vulkan/DX12/DX11 driver calls from the x64 binary reach the same native ARM64 GPU driver via interop thunks
- RenderDoc compatibility: RenderDoc only ships as x64. An x64 build is required for RenderDoc frame capture on WoA hardware. Measuring the emulation penalty tells us whether x64 benchmark results are still meaningful for comparison.
- Real-world relevance: Many Windows applications remain x64-only. Understanding the GPU benchmark impact of emulation helps assess whether WoA devices can be trusted for performance-sensitive x64 workloads.
- Isolating CPU vs GPU overhead: Since shaders run natively regardless of binary architecture, any performance difference is purely CPU-side (API call overhead, frame loop, buffer management). This cleanly separates translation overhead from GPU execution time.
| Metric | ARM64 Native | x64 Emulated | Delta |
|---|---|---|---|
| Vulkan | |||
| Avg FPS | 105.8 | 78.6 | −25.7% |
| Compute (ms) | 3.299 | 3.119 | −5.5% |
| Render (ms) | 5.704 | 5.586 | −2.1% |
| Total GPU (ms) | 9.002 | 8.706 | −3.3% |
| Frame Time (ms) | 9.448 | 12.726 | +34.7% |
| DX12 | |||
| Avg FPS | 116.7 | 98.3 | −15.8% |
| Compute (ms) | 3.281 | 3.257 | −0.7% |
| Render (ms) | 4.767 | 4.368 | −8.4% |
| Total GPU (ms) | 8.053 | 7.630 | −5.3% |
| Frame Time (ms) | 8.570 | 10.169 | +18.7% |
| DX11 | |||
| Avg FPS | 83.2 | 75.2 | −9.6% |
| Compute (ms) | 4.374 | 4.381 | +0.2% |
| Render (ms) | 4.579 | 4.470 | −2.4% |
| Total GPU (ms) | 11.775 | 11.678 | −0.8% |
| Frame Time (ms) | 12.024 | 13.293 | +10.6% |
GPU compute/render times are virtually identical between ARM64 and x64 builds — within ±5% across all APIs, well within run-to-run variance. This confirms the prediction: GPU shaders execute natively on the Adreno 640 regardless of the host binary's architecture. The translation layer does not affect GPU-side workload execution.
FPS is 10–26% lower on x64, with the penalty varying by API:
| API | FPS Penalty | Frame Time Increase | CPU Overhead Added |
|---|---|---|---|
| Vulkan | −25.7% | +3.28 ms | ~3.3 ms per frame |
| DX12 | −15.8% | +1.60 ms | ~1.6 ms per frame |
| DX11 | −9.6% | +1.27 ms | ~1.3 ms per frame |
Why Vulkan is penalised most heavily:
Vulkan has the highest per-frame CPU call count of the three APIs — explicit command buffer recording, descriptor set binding, fence management, and swapchain acquisition each require individual API calls that pass through the Prism translation layer. Each translated call adds a small overhead (~microseconds), but at 100+ calls per frame, the total accumulates to ~3.3 ms.
DX12 has slightly fewer per-frame CPU calls due to its command list model, resulting in a smaller 1.6 ms penalty. DX11's implicit driver handles most resource management internally (fewer API calls from the application), so it suffers the least translation overhead at 1.3 ms.
Why GPU times are slightly lower on x64 (counter-intuitive):
The x64 build's GPU compute and render times are marginally lower (by 1–5%) than ARM64. This is not because x64 code makes the GPU faster — it is an artefact of the higher frame time. With longer gaps between frame submissions (due to CPU translation overhead), the GPU has slightly more time to process each frame without contention, resulting in marginally cleaner timestamps. The difference is within measurement noise and should not be interpreted as a real GPU performance improvement.
-
GPU benchmark data from x64 builds is valid. GPU compute and render times are unaffected by Prism translation — x64 results can be directly compared against ARM64 results for GPU performance analysis.
-
FPS comparisons require a correction factor. x64 FPS is 10–26% lower than ARM64 native due to CPU-side translation overhead. When comparing Adreno 640 FPS against desktop GPUs (which run x64 natively), the ARM64 native results should be used as the true performance baseline.
-
Prism overhead is workload-dependent. API-heavy workloads (Vulkan, DX12) are penalised more than API-light workloads (DX11). For GPU-bound scenarios (this benchmark at 1M particles on Adreno 640), the penalty is modest because most frame time is spent on GPU execution, not CPU API calls.
-
RenderDoc x64 captures on WoA are viable. Since the x64 build produces identical GPU behaviour with only a CPU overhead penalty, RenderDoc frame captures from x64 builds are representative of true GPU workload behaviour — the captured GPU commands, timings, and resource state will match what the ARM64 native build would produce.







