forked from procrastineto/GPU-PCIe-Test
-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathmain_gui.cpp
More file actions
7940 lines (6886 loc) · 348 KB
/
Copy pathmain_gui.cpp
File metadata and controls
7940 lines (6886 loc) · 348 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
// ============================================================================
// GPU-PCIe-Test v3.4 - GUI Edition
// Dear ImGui + D3D12 Frontend
// ============================================================================
// Graphical frontend for the GPU/PCIe benchmark tool.
//
// UNIFIED BENCHMARK METHODOLOGY
// ─────────────────────────────────────────────────────────────────────────────
// This methodology is designed to be API-agnostic. The same approach applies
// to both the D3D12 and Vulkan versions, ensuring comparable results that can
// be cross-validated. The key principle: always test through the GPU's DMA
// copy engines directly, never through the graphics command processor.
//
// Queue Selection (the foundation of consistent results):
// D3D12 → D3D12_COMMAND_LIST_TYPE_DIRECT queue. The driver internally
// routes CopyResource calls to the GPU's DMA copy engines.
// (COPY queues are unreliable on some driver/hardware combos,
// notably eGPU over Thunderbolt, so DIRECT is used for safety.)
// Vulkan → Dedicated transfer queue family (VK_QUEUE_TRANSFER_BIT, no
// VK_QUEUE_GRAPHICS_BIT). Maps directly to DMA copy engines.
// Note: D3D12 DIRECT queue adds driver-level scheduling that may
// auto-optimize copy routing. Vulkan transfer queues give raw
// DMA engine access. Small measurement differences are expected.
//
// Bandwidth Tests:
// Download (GPU→CPU):
// GPU timestamps bracketing the copy commands on the DIRECT queue.
// D3D12: EndQuery(TIMESTAMP) before/after CopyResource.
//
// Upload (CPU→GPU) - Discrete GPUs:
// CPU round-trip timing. Records: upload + readback from same buffer.
// Uses previously measured download speed to subtract download time:
// upload_speed = data_size / (round_trip_time - download_time).
// Required because ReBAR/BAR mapping can cause GPU timestamps to report
// completion before data actually reaches VRAM over the PCIe bus.
//
// Upload (CPU→GPU) - Integrated GPUs:
// GPU timestamps, same as download. No ReBAR issue because both "CPU"
// and "GPU" memory are the same physical RAM with no bus transfer.
//
// Bidirectional Test:
// Dual DIRECT queues submitted simultaneously:
// Queue 1: upload copies (CPU→GPU)
// Queue 2: download copies (GPU→CPU)
// Wait for both fences, measure total wall-clock time.
// Total bandwidth = (upload_bytes + download_bytes) / elapsed_time.
// Falls back to single-queue interleaved copies if creation fails.
// Vulkan equivalent: 2 × dedicated transfer queues.
//
// Latency Tests:
// GPU timestamps per individual small copy on the DIRECT queue.
// D3D12: EndQuery(TIMESTAMP) which serializes with prior work.
//
// Command Latency:
// Back-to-back timestamp pairs with no work between them.
// Measures minimum per-command dispatch overhead of the DIRECT queue.
//
// CROSS-API EQUIVALENCES
// ─────────────────────────────────────────────────────────────────────────────
// D3D12 Vulkan
// ──────────────────────────────── ──────────────────────────────────────
// ID3D12Device (bench) VkDevice (bench)
// COMMAND_LIST_TYPE_DIRECT queue Transfer queue family
// (driver routes to DMA engine) (direct DMA engine access)
// CopyResource / CopyBufferRegion vkCmdCopyBuffer
// EndQuery(TIMESTAMP) vkCmdWriteTimestamp
// ID3D12QueryHeap (TIMESTAMP) VkQueryPool (TIMESTAMP)
// ResolveQueryData + Map readback vkGetQueryPoolResults
// GetTimestampFrequency (ticks/sec) benchTimestampPeriod (ns/tick)
// D3D12_HEAP_TYPE_UPLOAD VK_MEMORY_PROPERTY_HOST_VISIBLE
// D3D12_HEAP_TYPE_DEFAULT VK_MEMORY_PROPERTY_DEVICE_LOCAL
// D3D12_HEAP_TYPE_READBACK HOST_VISIBLE + HOST_CACHED
// 2 × DIRECT queues (bidir) Dual transfer queues (bidir)
//
// ============================================================================
// Features:
// - Real-time progress visualization
// - Interactive configuration
// - Results graphs and charts with standard comparisons
// - CSV export
// - VRAM-aware buffer sizing
// - VRAM integrity scanning (multiple test patterns, error clustering)
// - eGPU auto-detection (Thunderbolt/USB4/USB via device tree)
// - Integrated GPU (APU) proper detection - no fake PCIe reporting
// - Actual PCIe link detection via SetupAPI
// - Improved upload measurement: uses measured download speed for accuracy
// - System RAM detection via WMI (speed, channels, type)
// - Full Unicode support (UTF-8 internal)
// ============================================================================
// Uncomment to enable debug logging for external GPU detection
// This will log the device tree traversal to help diagnose detection issues
// #define DEBUG_EXTERNAL_DETECTION
#ifndef NOMINMAX
#define NOMINMAX
#endif
#define _WIN32_DCOM // Required for CoInitializeEx
#include <windows.h>
#include <commdlg.h>
#include <shlobj.h>
#include <d3d12.h>
#include <d3dcompiler.h>
#include <dxgi1_6.h>
#include <wrl.h>
#include <setupapi.h>
#include <devpkey.h>
#include <cfgmgr32.h>
#include <initguid.h>
#include <devpropdef.h>
#include <comdef.h>
#include <wbemidl.h>
#include <iostream>
#include <vector>
#include <string>
#include <algorithm>
#include <numeric>
#include <chrono>
#include <ctime>
#include <thread>
#include <atomic>
#include <mutex>
#include <fstream>
#include <sstream>
#include <iomanip>
#include <cmath>
#include <map>
#include <set>
#include <random>
#include <array>
#pragma comment(lib, "d3d12.lib")
#pragma comment(lib, "dxgi.lib")
#pragma comment(lib, "setupapi.lib")
#pragma comment(lib, "cfgmgr32.lib")
#pragma comment(lib, "wbemuuid.lib")
#pragma comment(lib, "ole32.lib")
#pragma comment(lib, "oleaut32.lib")
#pragma comment(lib, "d3dcompiler.lib")
// ImGui headers (must be in imgui/ subfolder)
#include "imgui/imgui.h"
#include "imgui/imgui_impl_win32.h"
#include "imgui/imgui_impl_dx12.h"
#include "imgui/implot.h"
#include "imgui/imgui_internal.h"
using Microsoft::WRL::ComPtr;
// ============================================================================
// CONSTANTS
// ============================================================================
namespace Constants {
constexpr int LATENCY_WARMUP_ITERATIONS = 100;
constexpr int WINDOW_WIDTH = 1400;
constexpr int WINDOW_HEIGHT = 900;
constexpr int NUM_FRAMES_IN_FLIGHT = 3;
constexpr size_t DEFAULT_BANDWIDTH_SIZE = 256ull * 1024 * 1024;
constexpr size_t DEFAULT_LATENCY_SIZE = 1;
constexpr int DEFAULT_BANDWIDTH_BATCHES = 32;
constexpr int DEFAULT_COPIES_PER_BATCH = 8;
constexpr int DEFAULT_LATENCY_ITERS = 2000; // Reduced from 10000 - still accurate, much faster
constexpr int DEFAULT_NUM_RUNS = 3;
constexpr float BASE_FONT_SCALE = 1.0f; // User's preferred 2x scaling
// Timeout and retry constants
constexpr DWORD FENCE_WAIT_TIMEOUT_MS = 8000; // 8 seconds per fence wait
constexpr int MAX_FENCE_RETRIES = 3; // Max retries before aborting
constexpr DWORD GLOBAL_BENCHMARK_TIMEOUT_MS = 300000; // 5 minute global timeout
// VRAM safety margin (leave 20% free for system use)
constexpr double VRAM_SAFETY_MARGIN = 0.8;
constexpr size_t MIN_BANDWIDTH_SIZE = 16ull * 1024 * 1024; // 16 MB minimum
// Bandwidth is reported in decimal GB/s (1 GB = 1e9 bytes) to match how
// PCIe/TB/USB4 interface standards are specified. (Was 1024^3, which
// under-reported by ~7.4% against every standard in the table.)
constexpr double BYTES_PER_GB = 1e9;
// eGPU detection thresholds
constexpr double EGPU_BANDWIDTH_THRESHOLD = 5.0; // GB/s - below this suggests external connection
constexpr double TB3_MAX_BANDWIDTH = 3.5; // Thunderbolt 3 typical max
constexpr double TB4_MAX_BANDWIDTH = 4.5; // Thunderbolt 4 / USB4 40Gbps typical max (32 Gbps PCIe tunnel)
constexpr double TB5_MAX_BANDWIDTH = 7.5; // Thunderbolt 5 / USB4 80Gbps typical max (64 Gbps PCIe tunnel)
// Log and UI limits
constexpr int MAX_LOG_LINES = 500;
// VRAM scanning chunk sizes
constexpr size_t VRAM_CHUNK_PREFERRED = 512ull * 1024 * 1024; // 512 MB per chunk
constexpr size_t VRAM_CHUNK_MINIMUM = 128ull * 1024 * 1024; // 128 MB minimum viable
// Maximum bandwidth test buffer (2 GB)
constexpr size_t MAX_BANDWIDTH_BUFFER = 2ull * 1024 * 1024 * 1024;
// eGPU device tree traversal depth
constexpr int EGPU_MAX_TREE_DEPTH = 15;
// Inter-batch sleep to avoid starving other work (microseconds)
constexpr int BENCHMARK_SLEEP_US = 50;
// Memory latency compute shader test
constexpr size_t MEMORY_LATENCY_BUFFER_SIZE = 32ull * 1024 * 1024; // 32 MB (> typical GPU L2 cache)
constexpr uint32_t MEMORY_LATENCY_NUM_CHASES = 100000; // Pointer chases per dispatch
constexpr int MEMORY_LATENCY_WARMUP_DISPATCHES = 3; // Warmup dispatches (stabilize clocks)
constexpr int MEMORY_LATENCY_MEASURE_DISPATCHES = 10; // Measurement dispatches
}
// Embedded HLSL compute shader for GPU memory latency measurement (pointer-chase)
// A single thread chases a randomly-shuffled linked list through a large buffer,
// forcing dependent loads that cannot be prefetched. Each load measures true memory
// round-trip latency from the GPU's perspective.
static const char* g_memoryLatencyHLSL = R"(
RWStructuredBuffer<uint> chain : register(u0);
cbuffer Params : register(b0) {
uint numChases;
uint startIndex;
};
[numthreads(1, 1, 1)]
void CSMain(uint3 id : SV_DispatchThreadID) {
uint idx = startIndex;
[loop]
for (uint i = 0; i < numChases; i++) {
idx = chain[idx];
}
chain[0] = idx; // prevent dead code elimination
}
)";
// ============================================================================
// DATA STRUCTURES
// ============================================================================
struct GPUInfo {
std::string name;
std::string vendor;
uint32_t vendorId = 0;
uint32_t deviceId = 0;
size_t dedicatedVRAM = 0;
size_t sharedMemory = 0;
bool isIntegrated = false;
bool isValid = true; // False for "no GPU found" placeholder
// PCIe link information (detected via SetupAPI)
int pcieGenCurrent = 0; // Current PCIe generation (1-6)
int pcieLanesCurrent = 0; // Current lane width (1,2,4,8,16,32)
int pcieGenMax = 0; // Max supported PCIe generation
int pcieLanesMax = 0; // Max supported lane width
bool pcieInfoValid = false; // True if we successfully queried PCIe info
std::string pcieLocationPath; // Device location path for identification
// Thunderbolt/USB4/USB detection (detected via device tree topology)
bool isThunderbolt = false; // Connected via Thunderbolt (Intel certified)
bool isUSB4 = false; // Connected via USB4 (includes TB4, AMD USB4)
bool isUSB = false; // Connected via any USB (USB3/USB4/TB)
int thunderboltVersion = 0; // 3 or 4 (0 if unknown or not TB)
std::string externalConnectionType; // "Thunderbolt 4", "USB4", "AMD USB4", etc.
// Store adapter for reliable selection (handles hot-plug scenarios)
ComPtr<IDXGIAdapter1> adapter;
};
// System RAM information (detected via WMI)
struct SystemMemoryInfo {
uint32_t speedMT = 0; // Speed in MT/s (e.g., 6400)
uint32_t configuredSpeedMT = 0; // Configured/actual speed in MT/s
uint32_t channels = 0; // Number of channels (1, 2, 4, etc.)
uint32_t totalSticks = 0; // Number of physical DIMMs
uint64_t totalCapacityGB = 0; // Total capacity in GB
std::string type; // "DDR4", "DDR5", "LPDDR5", etc.
std::string formFactor; // "DIMM", "SODIMM", etc.
double theoreticalBandwidth = 0; // Calculated theoretical bandwidth in GB/s
bool detected = false; // True if WMI query succeeded
std::string errorMessage; // Error message if detection failed
double ratedLatencyNs = 0; // Estimated chip latency from speed tier (ns)
int estimatedCL = 0; // Estimated CAS latency
bool latencyEstimated = false; // True if we found a matching speed tier
uint32_t busWidthBits = 0; // Memory bus width from SMBIOS data widths (0 = unknown)
bool channelsUnverified = false; // Soldered LPDDR: channel count is only an SMBIOS-based guess
bool channelsInferred = false; // Channel count raised to fit the measured iGPU bandwidth
bool needsElevation = false; // Linux: SMBIOS (dmidecode) needs root; per-device details missing
};
// VRAM test pattern types
enum class VRAMTestPattern {
AllZeros, // 0x00000000
AllOnes, // 0xFFFFFFFF
Checkerboard, // 0xAAAAAAAA
InverseCheckerboard,// 0x55555555
Random, // Random data
MarchingOnes, // Walking 1 pattern
MarchingZeros, // Walking 0 pattern
AddressPattern // Address-based pattern for location detection
};
// Classification of a memory error by physical mechanism. Knowing the kind
// of error helps diagnose whether the fault is in a single wire, a chip's
// internal logic, or the address bus.
enum class VRAMErrorKind {
SingleBit, // Exactly one bit flipped (single-wire / single-cell error)
MultiBit, // 2-6 bits flipped (multi-bit transmission/burst error)
AddressBus, // 7+ bits flipped with random distribution (wrong cell returned)
StuckAtZero, // Read back 0x00000000 when non-zero was written
StuckAtOne, // Read back 0xFFFFFFFF when non-FF was written
RefreshError // Mismatch appeared on re-read after initial write succeeded
};
inline const char* GetErrorKindName(VRAMErrorKind k) {
switch (k) {
case VRAMErrorKind::SingleBit: return "single-bit";
case VRAMErrorKind::MultiBit: return "multi-bit";
case VRAMErrorKind::AddressBus: return "address-bus";
case VRAMErrorKind::StuckAtZero: return "stuck-at-0";
case VRAMErrorKind::StuckAtOne: return "stuck-at-1";
case VRAMErrorKind::RefreshError: return "refresh";
default: return "unknown";
}
}
// VRAM error information
struct VRAMError {
size_t offsetStart = 0; // Start offset of error region
size_t offsetEnd = 0; // End offset of error region
uint32_t expected = 0; // Expected value
uint32_t actual = 0; // Actual value read back
VRAMTestPattern pattern; // Which pattern detected the error
size_t errorCount = 0; // Number of errors in this region
// Bit-level diagnostic info (populated by CompareBuffers for the first error in the cluster)
uint32_t bitFlipMask = 0; // XOR(expected, actual) - bits that flipped
uint32_t bitFlipCount = 0; // popcount(bitFlipMask)
int bitIndex = -1; // For SingleBit errors: 0-31; otherwise -1
VRAMErrorKind kind = VRAMErrorKind::SingleBit;
};
// VRAM test results
struct VRAMTestResult {
bool completed = false;
bool cancelled = false;
size_t totalBytesTested = 0;
size_t totalErrors = 0;
std::vector<VRAMError> errors;
std::vector<std::string> patternResults; // Results per pattern
double testDurationSeconds = 0;
std::string summary;
// Aggregate bit-level error stats (across ALL chunks and patterns):
std::array<size_t, 32> bitFlipHistogram = {}; // bitFlipHistogram[b] = times bit b flipped
std::array<size_t, 6> errorKindCounts = {}; // Indexed by VRAMErrorKind enum order
size_t refreshPassErrors = 0; // Errors detected during re-read phase only
};
// PCIe generation speed in GT/s (gigatransfers per second)
static const double PCIE_GEN_SPEEDS[] = {
0.0, // Gen 0 (invalid)
2.5, // Gen 1
5.0, // Gen 2
8.0, // Gen 3
16.0, // Gen 4
32.0, // Gen 5
64.0 // Gen 6
};
// Protocol overhead efficiency factors (accounts for TLP headers, DLLPs, flow control)
// Gen 1-2: 8b/10b encoding (80%) + ~15% protocol overhead = ~65-70% efficiency
// Gen 3+: 128b/130b encoding (98.5%) + ~10-15% protocol overhead = ~85% efficiency
static const double PCIE_PROTOCOL_EFFICIENCY[] = {
0.0, // Gen 0 (invalid)
0.65, // Gen 1 (8b/10b + high overhead)
0.70, // Gen 2 (8b/10b + moderate overhead)
0.85, // Gen 3 (128b/130b + typical overhead)
0.85, // Gen 4 (128b/130b + typical overhead)
0.85, // Gen 5 (128b/130b + typical overhead)
0.85 // Gen 6 (128b/130b + typical overhead)
};
// Calculate theoretical (raw) bandwidth for PCIe config (in GB/s)
// This is the maximum possible with perfect efficiency (encoding overhead only)
inline double CalculatePCIeBandwidth(int gen, int lanes) {
if (gen < 1 || gen > 6 || lanes < 1) return 0.0;
// PCIe uses 128b/130b encoding for Gen3+, 8b/10b for Gen1-2
double encodingEfficiency = (gen >= 3) ? (128.0 / 130.0) : (8.0 / 10.0);
double gtPerSec = PCIE_GEN_SPEEDS[gen];
// GT/s * lanes * encoding efficiency / 8 bits per byte = GB/s
return (gtPerSec * lanes * encodingEfficiency) / 8.0;
}
// Calculate realistic (achievable) bandwidth accounting for protocol overhead
// This is what well-optimized software can typically achieve
inline double CalculateRealisticPCIeBandwidth(int gen, int lanes) {
if (gen < 1 || gen > 6 || lanes < 1) return 0.0;
double gtPerSec = PCIE_GEN_SPEEDS[gen];
double efficiency = PCIE_PROTOCOL_EFFICIENCY[gen];
// GT/s * lanes * overall efficiency / 8 bits per byte = GB/s
return (gtPerSec * lanes * efficiency) / 8.0;
}
struct BenchmarkResult {
std::string testName;
double minValue = 0;
double avgValue = 0;
double maxValue = 0;
std::string unit;
std::vector<double> samples; // For graphing
};
// VRAM scan preset modes. Selecting a preset snaps all individual VRAM scan
// options to predefined values; editing any individual option switches the
// preset to Custom.
enum class VRAMScanPreset {
Quick, // 4 patterns, no re-read, no non-sequential, 50% coverage (~30s on 8GB)
Standard, // All 8 patterns, no re-read, no non-sequential, 80% coverage (~2-3min) - default
Deep, // All 8 + re-read x4 + non-sequential 64KB + pre-heat 30s + GPU verify, 90% (~5min)
Thorough, // All 8 + re-read x10 + non-sequential 64KB + pre-heat 60s + GPU verify, 95% (~12-15min)
Marathon, // Same as Thorough but loops indefinitely until cancelled
Custom // User has manually edited individual options
};
struct BenchmarkConfig {
size_t bandwidthSize = Constants::DEFAULT_BANDWIDTH_SIZE;
size_t latencySize = Constants::DEFAULT_LATENCY_SIZE;
int bandwidthBatches = Constants::DEFAULT_BANDWIDTH_BATCHES;
int copiesPerBatch = Constants::DEFAULT_COPIES_PER_BATCH;
int latencyIters = Constants::DEFAULT_LATENCY_ITERS;
int numRuns = Constants::DEFAULT_NUM_RUNS;
bool runBidirectional = true;
bool runLatency = true;
bool runMemoryLatency = true; // GPU memory latency via compute shader pointer-chase
bool quickMode = false;
bool averageRuns = true; // When false, record each run individually
bool debugLogging = false; // Verbose diagnostic logging for memory latency test etc.
int selectedGPU = 0;
// VRAM scan options (borrowed from memtest_vulkan-style stress testing)
// Pattern order matches VRAMTestPattern enum:
// 0=AllZeros, 1=AllOnes, 2=Checkerboard, 3=InverseCheckerboard,
// 4=Random, 5=MarchingOnes, 6=MarchingZeros, 7=AddressPattern
VRAMScanPreset vramScanPreset = VRAMScanPreset::Standard;
std::array<bool, 8> vramPatternsEnabled = { true, true, true, true, true, true, true, true };
bool vramRereadEnabled = false; // Re-read pass to catch refresh/retention errors
int vramRereadIterations = 4; // Number of re-reads (1-20)
bool vramNonSequentialEnabled = false; // Shuffle block read order to defeat row buffer caching
int vramNonSequentialBlockSize = 65536; // 16384 / 65536 / 262144 (16/64/256 KB)
int vramCoveragePercent = 80; // 50 / 80 / 90 / 95
bool vramPreheatEnabled = false; // Run GPU heat-up load before scan to catch thermal errors
int vramPreheatSeconds = 30; // Pre-heat duration (10-120s)
bool vramMarathonMode = false; // Loop forever until cancelled (set by Marathon preset)
bool vramGpuVerify = false; // Use GPU compute shader for verification instead of CPU readback
};
struct InterfaceSpeed {
const char* name;
double bandwidth; // Realistic achievable bandwidth (with protocol overhead)
double theoretical; // Raw theoretical bandwidth (encoding overhead only)
const char* description;
bool tunneledExternal; // true = TB/USB4 PCIe-tunneling tier (eGPU over TB/USB4)
};
// Updated interface standards with realistic achievable bandwidth
// PCIe Gen3+ uses ~85% efficiency (128b/130b encoding + TLP/DLLP overhead)
// Thunderbolt/USB4 PCIe payload is capped by the PCIe *tunnel*, not the link
// rate: 40 Gbps links tunnel at most 32 Gbps of PCIe (4.0 GB/s raw), 80 Gbps
// links (TB5 / USB4v2) tunnel at most 64 Gbps (8.0 GB/s raw). TB5 and USB4
// 80Gbps are indistinguishable from bandwidth alone, so they share one entry.
static const InterfaceSpeed INTERFACE_SPEEDS[] = {
{"PCIe 3.0 x4", 3.40, 3.94, "Entry-level GPU slot", false},
{"PCIe 3.0 x8", 6.80, 7.88, "Mid-range GPU slot", false},
{"PCIe 3.0 x16", 13.60, 15.75, "Standard discrete GPU", false},
{"PCIe 4.0 x4", 6.80, 7.88, "NVMe / Entry eGPU", false},
{"PCIe 4.0 x8", 13.60, 15.75, "Mid-range PCIe 4.0", false},
{"PCIe 4.0 x16", 27.20, 31.51, "High-end discrete GPU", false},
{"PCIe 5.0 x8", 27.20, 31.51, "PCIe 5.0 mid-range", false},
{"PCIe 5.0 x16", 54.40, 63.02, "High-end PCIe 5.0 GPU slot", false},
{"PCIe 6.0 x16", 108.80, 126.03, "Next-gen PCIe 6.0 GPU slot", false},
{"OCuLink 1.0", 3.40, 3.94, "PCIe 3.0 x4 external", false},
{"OCuLink 2.0", 6.80, 7.88, "PCIe 4.0 x4 external", false},
{"Thunderbolt 3", 2.50, 2.80, "40 Gbps link (variable PCIe allocation)", true},
{"Thunderbolt 4 / USB4 40Gbps", 3.50, 4.00, "40 Gbps link (32 Gbps PCIe tunnel)", true},
{"Thunderbolt 5 / USB4 80Gbps", 6.50, 8.00, "80 Gbps link (64 Gbps PCIe tunnel)", true},
};
static const int NUM_INTERFACE_SPEEDS = sizeof(INTERFACE_SPEEDS) / sizeof(INTERFACE_SPEEDS[0]);
// Memory bandwidth standards for integrated GPU (APU) comparison
// Realistic = ~80% of theoretical (memory controller overhead, contention)
// Theoretical = speed_MT/s * 8 bytes * 2 channels / 1000
static const InterfaceSpeed MEMORY_STANDARDS[] = {
{"DDR4-2400 DC", 30.7, 38.4, "Dual-channel DDR4-2400"},
{"DDR4-3200 DC", 41.0, 51.2, "Dual-channel DDR4-3200"},
{"DDR5-4800 DC", 61.4, 76.8, "Dual-channel DDR5-4800"},
{"DDR5-5600 DC", 71.7, 89.6, "Dual-channel DDR5-5600"},
{"DDR5-6400 DC", 81.9, 102.4, "Dual-channel DDR5-6400"},
{"DDR5-8800 DC", 112.6, 140.8, "Dual-channel DDR5-8800"},
{"LPDDR5X-7500", 96.0, 120.0, "Dual-channel LPDDR5X-7500"},
{"LPDDR5X-8533", 109.2, 136.5, "Dual-channel LPDDR5X-8533"},
{"DDR6-12800 DC", 163.8, 204.8, "Dual-channel DDR6-12800 (projected)"},
{"DDR6-17600 DC", 225.3, 281.6, "Dual-channel DDR6-17600 (projected)"},
};
static const int NUM_MEMORY_STANDARDS = sizeof(MEMORY_STANDARDS) / sizeof(MEMORY_STANDARDS[0]);
// Typical CAS latency by speed tier for rated chip latency estimation
// Formula: latency_ns = CL / (speedMT / 2000.0) * 1000.0
struct MemoryLatencyEntry {
uint32_t speedMT; // Speed in MT/s
const char* type; // "DDR4", "DDR5", "DDR6", etc.
int typicalCL; // Typical CAS latency for this speed tier
double latencyNs; // Pre-calculated chip latency in nanoseconds
};
static const MemoryLatencyEntry MEMORY_LATENCY_TABLE[] = {
{ 2400, "DDR4", 17, 14.17 },
{ 2666, "DDR4", 17, 12.76 },
{ 3200, "DDR4", 16, 10.00 },
{ 3600, "DDR4", 18, 10.00 },
{ 4800, "DDR5", 40, 16.67 },
{ 5200, "DDR5", 38, 14.62 },
{ 5600, "DDR5", 36, 12.86 },
{ 6000, "DDR5", 36, 12.00 },
{ 6400, "DDR5", 40, 12.50 },
{ 6800, "DDR5", 40, 11.76 },
{ 7200, "DDR5", 40, 11.11 },
{ 7500, "LPDDR5X", 36, 9.60 },
{ 7600, "DDR5", 40, 10.53 },
{ 8000, "DDR5", 40, 10.00 },
{ 8533, "LPDDR5X", 36, 8.44 },
{ 8800, "DDR5", 44, 10.00 },
{ 12800, "DDR6", 52, 8.13 }, // projected
{ 17600, "DDR6", 60, 6.82 }, // projected
};
static const int NUM_MEMORY_LATENCY_ENTRIES = sizeof(MEMORY_LATENCY_TABLE) / sizeof(MEMORY_LATENCY_TABLE[0]);
// ============================================================================
// APPLICATION STATE
// ============================================================================
enum class AppState { Idle, Running, Completed };
// Fence wait result for robust error handling
enum class FenceWaitResult { Success, Timeout, Error, Cancelled };
struct AppContext {
// Window
HWND hwnd = nullptr;
int windowWidth = Constants::WINDOW_WIDTH;
int windowHeight = Constants::WINDOW_HEIGHT;
// D3D12 Device (for rendering)
ComPtr<ID3D12Device> device;
ComPtr<ID3D12CommandQueue> commandQueue;
ComPtr<IDXGISwapChain3> swapChain;
ComPtr<ID3D12DescriptorHeap> rtvHeap;
ComPtr<ID3D12DescriptorHeap> srvHeap;
ComPtr<ID3D12CommandAllocator> commandAllocators[Constants::NUM_FRAMES_IN_FLIGHT];
ComPtr<ID3D12GraphicsCommandList> commandList;
ComPtr<ID3D12Resource> renderTargets[Constants::NUM_FRAMES_IN_FLIGHT];
ComPtr<ID3D12Fence> fence;
HANDLE fenceEvent = nullptr;
UINT64 fenceValues[Constants::NUM_FRAMES_IN_FLIGHT] = {};
UINT64 currentFenceValue = 0;
UINT frameIndex = 0;
UINT rtvDescriptorSize = 0;
// D3D12 Benchmark Device (separate for benchmarking)
ComPtr<ID3D12Device> benchDevice;
ComPtr<ID3D12CommandQueue> benchQueue;
ComPtr<ID3D12CommandAllocator> benchAllocator;
ComPtr<ID3D12GraphicsCommandList> benchList;
ComPtr<ID3D12Fence> benchFence;
HANDLE benchFenceEvent = nullptr;
UINT64 benchFenceValue = 1;
D3D12_COMMAND_LIST_TYPE benchQueueType = D3D12_COMMAND_LIST_TYPE_DIRECT;
// Second queue for bidirectional transfers (allows true simultaneous upload/download)
ComPtr<ID3D12CommandQueue> benchQueue2;
ComPtr<ID3D12CommandAllocator> benchAllocator2;
ComPtr<ID3D12GraphicsCommandList> benchList2;
ComPtr<ID3D12Fence> benchFence2;
HANDLE benchFenceEvent2 = nullptr;
UINT64 benchFenceValue2 = 1;
bool hasDualQueues = false;
// GPU list
std::vector<GPUInfo> gpuList;
std::vector<std::string> gpuComboNames;
std::vector<const char*> gpuComboPointers;
// Config
BenchmarkConfig config;
// State
AppState state = AppState::Idle;
std::atomic<float> progress{ 0.0f };
std::atomic<float> overallProgress{ 0.0f };
std::atomic<int> currentRun{ 0 };
std::atomic<int> totalTests{ 0 };
std::atomic<int> completedTests{ 0 };
std::atomic<bool> cancelRequested{ false };
std::atomic<bool> benchmarkAborted{ false }; // For critical failures
std::atomic<int> fenceTimeoutCount{ 0 }; // Track consecutive timeouts
std::string currentTest;
std::mutex resultsMutex;
std::vector<BenchmarkResult> results;
std::thread benchmarkThread;
std::atomic<bool> benchmarkThreadRunning{ false }; // Track if thread is active
// Benchmark timing for global timeout
std::chrono::steady_clock::time_point benchmarkStartTime;
// UI State (removed g_ prefix - cleaner inside struct)
bool showResultsWindow = false;
bool showGraphsWindow = false;
bool showCompareWindow = false; // NEW: Compare to standards window
bool showAboutDialog = false;
bool dockingInitialized = false;
bool isResizing = false; // Was: g_isResizing
bool pendingResize = false; // Was: g_pendingResize
int pendingWidth = 0; // Was: g_pendingWidth
int pendingHeight = 0; // Was: g_pendingHeight
// Log buffer
std::mutex logMutex;
std::vector<std::string> logLines;
// Detected interface results
std::string detectedInterface;
std::string detectedInterfaceDescription;
double uploadBW = 0;
double downloadBW = 0;
double uploadPercentage = 0; // Percentage of closest standard
double downloadPercentage = 0; // Percentage of closest standard
std::string closestUploadStandard; // Name of closest standard
std::string closestDownloadStandard; // Name of closest standard
// eGPU detection
bool possibleEGPU = false;
std::string eGPUConnectionType;
// Integrated GPU memory info
std::string integratedMemoryType; // e.g., "DDR5"
std::string integratedFabricType; // e.g., "AMD Infinity Fabric"
// Summary window
bool showSummaryWindow = false;
double actualPCIeBandwidth = 0; // Theoretical max based on detected link
std::string actualPCIeConfig; // e.g., "PCIe 4.0 x16"
std::string summaryExplanation; // Explanation of measured vs actual
// System memory info (detected via WMI)
SystemMemoryInfo systemMemory;
// VRAM test state
std::atomic<bool> vramTestRunning{ false };
std::atomic<bool> vramTestCancelRequested{ false };
std::thread vramTestThread;
VRAMTestResult vramTestResult;
std::atomic<float> vramTestProgress{ 0.0f };
std::string vramTestCurrentPattern;
bool showVRAMTestWindow = false;
// (Coverage is now driven by g_app.config.vramCoveragePercent - see BenchmarkConfig)
};
static AppContext g_app;
// Helper to add log messages
void Log(const std::string& msg) {
std::lock_guard<std::mutex> lock(g_app.logMutex);
g_app.logLines.push_back(msg);
// Keep last N lines
if (g_app.logLines.size() > static_cast<size_t>(Constants::MAX_LOG_LINES)) {
g_app.logLines.erase(g_app.logLines.begin());
}
}
void ClearLog() {
std::lock_guard<std::mutex> lock(g_app.logMutex);
g_app.logLines.clear();
}
// currentTest and vramTestCurrentPattern are written by the benchmark/scan
// worker threads and read by the UI thread. Route every access through these
// helpers so the UI never observes a torn std::string (reallocation mid-read).
static std::mutex g_statusMutex;
void SetCurrentTest(const std::string& s) {
std::lock_guard<std::mutex> lock(g_statusMutex);
g_app.currentTest = s;
}
std::string GetCurrentTest() {
std::lock_guard<std::mutex> lock(g_statusMutex);
return g_app.currentTest;
}
void SetVramPattern(const std::string& s) {
std::lock_guard<std::mutex> lock(g_statusMutex);
g_app.vramTestCurrentPattern = s;
}
std::string GetVramPattern() {
std::lock_guard<std::mutex> lock(g_statusMutex);
return g_app.vramTestCurrentPattern;
}
// ============================================================================
// UNICODE HELPER FUNCTIONS (Full UTF-8 Support)
// ============================================================================
// Convert wide string (Windows native) to UTF-8 string
std::string WideToUtf8(const std::wstring& wide) {
if (wide.empty()) return std::string();
int size = WideCharToMultiByte(CP_UTF8, 0, wide.c_str(), static_cast<int>(wide.size()),
nullptr, 0, nullptr, nullptr);
if (size <= 0) return std::string();
std::string utf8(size, 0);
WideCharToMultiByte(CP_UTF8, 0, wide.c_str(), static_cast<int>(wide.size()),
&utf8[0], size, nullptr, nullptr);
return utf8;
}
// Convert UTF-8 string to wide string (Windows native)
std::wstring Utf8ToWide(const std::string& utf8) {
if (utf8.empty()) return std::wstring();
int size = MultiByteToWideChar(CP_UTF8, 0, utf8.c_str(), static_cast<int>(utf8.size()),
nullptr, 0);
if (size <= 0) return std::wstring();
std::wstring wide(size, 0);
MultiByteToWideChar(CP_UTF8, 0, utf8.c_str(), static_cast<int>(utf8.size()),
&wide[0], size);
return wide;
}
// Convert BSTR (COM) to UTF-8 string
std::string BstrToUtf8(BSTR bstr) {
if (!bstr) return std::string();
int len = SysStringLen(bstr);
if (len == 0) return std::string();
return WideToUtf8(std::wstring(bstr, len));
}
// ============================================================================
// SYSTEM MEMORY DETECTION (WMI)
// ============================================================================
// DDR type mapping from SMBIOSMemoryType
std::string GetDDRTypeFromSMBIOS(uint16_t memoryType) {
switch (memoryType) {
case 20: return "DDR";
case 21: return "DDR2";
case 22: return "DDR2 FB-DIMM";
case 24: return "DDR3";
case 26: return "DDR4";
case 27: return "LPDDR";
case 28: return "LPDDR2";
case 29: return "LPDDR3";
case 30: return "LPDDR4";
case 34: return "DDR5";
case 35: return "LPDDR5";
case 36: return "LPDDR5X";
// DDR6/DDR7: JEDEC SMBIOS type codes not yet assigned as of 2025.
// When assigned, add them here. Expected in SMBIOS 3.8+.
// case ??: return "DDR6";
// case ??: return "LPDDR6";
// case ??: return "DDR7";
default: return "Unknown";
}
}
// Soldered memory (LPDDR on laptops / APUs; SMBIOS form factor "Row Of Chips")
// has no DIMM-per-channel relationship: firmware may describe a 256-bit bus as
// one device or as eight 32-bit devices. Device counting and locator letters
// therefore say little about the channel count.
static bool IsSolderedMemory(const std::string& type, const std::string& formFactor) {
auto lower = [](std::string s) {
std::transform(s.begin(), s.end(), s.begin(),
[](char c) { return static_cast<char>(::tolower(static_cast<unsigned char>(c))); });
return s;
};
std::string t = lower(type), f = lower(formFactor);
return t.find("lpddr") != std::string::npos ||
f.find("row of chips") != std::string::npos ||
f.find("chip") != std::string::npos;
}
// Form factor mapping
std::string GetFormFactorName(uint16_t formFactor) {
switch (formFactor) {
case 8: return "DIMM";
case 9: return "Row Of Chips"; // soldered (LPDDR on laptops / APUs)
case 12: return "SODIMM";
case 13: return "SRIMM";
case 14: return "FB-DIMM";
default: return "Unknown";
}
}
// Detect system memory configuration via WMI
SystemMemoryInfo DetectSystemMemory() {
SystemMemoryInfo info;
HRESULT hr = CoInitializeEx(0, COINIT_MULTITHREADED);
// Track if WE initialized COM (S_OK or S_FALSE means we added a reference)
// RPC_E_CHANGED_MODE means COM was already init'd with different mode - we didn't init
bool weInitializedCom = (hr == S_OK || hr == S_FALSE);
bool comAvailable = SUCCEEDED(hr) || hr == RPC_E_CHANGED_MODE;
if (!comAvailable) {
info.errorMessage = "Failed to initialize COM";
return info;
}
// Set security levels
hr = CoInitializeSecurity(
nullptr, -1, nullptr, nullptr,
RPC_C_AUTHN_LEVEL_DEFAULT,
RPC_C_IMP_LEVEL_IMPERSONATE,
nullptr, EOAC_NONE, nullptr
);
// Ignore security errors if already initialized
IWbemLocator* pLoc = nullptr;
hr = CoCreateInstance(CLSID_WbemLocator, 0, CLSCTX_INPROC_SERVER,
IID_IWbemLocator, reinterpret_cast<void**>(&pLoc));
if (FAILED(hr)) {
info.errorMessage = "Failed to create WbemLocator";
if (weInitializedCom) CoUninitialize();
return info;
}
IWbemServices* pSvc = nullptr;
hr = pLoc->ConnectServer(
_bstr_t(L"ROOT\\CIMV2"), nullptr, nullptr, 0, 0, 0, 0, &pSvc
);
if (FAILED(hr)) {
info.errorMessage = "Failed to connect to WMI";
pLoc->Release();
if (weInitializedCom) CoUninitialize();
return info;
}
// Set proxy security
hr = CoSetProxyBlanket(
pSvc, RPC_C_AUTHN_WINNT, RPC_C_AUTHZ_NONE, nullptr,
RPC_C_AUTHN_LEVEL_CALL, RPC_C_IMP_LEVEL_IMPERSONATE, nullptr, EOAC_NONE
);
// Query physical memory
IEnumWbemClassObject* pEnumerator = nullptr;
hr = pSvc->ExecQuery(
_bstr_t(L"WQL"),
_bstr_t(L"SELECT * FROM Win32_PhysicalMemory"),
WBEM_FLAG_FORWARD_ONLY | WBEM_FLAG_RETURN_IMMEDIATELY,
nullptr, &pEnumerator
);
if (FAILED(hr)) {
info.errorMessage = "WMI query failed";
pSvc->Release();
pLoc->Release();
if (weInitializedCom) CoUninitialize();
return info;
}
// Process results
IWbemClassObject* pclsObj = nullptr;
ULONG uReturn = 0;
uint32_t maxSpeed = 0;
uint32_t maxConfiguredSpeed = 0;
uint64_t totalCapacity = 0;
std::string detectedType;
std::string detectedFormFactor;
uint16_t detectedMemoryType = 0; // SMBIOS memory type for DDR generation detection
int stickCount = 0;
uint32_t dataWidthSum = 0; // sum of per-device SMBIOS data widths (bits), populated devices only
std::set<std::string> uniqueBanks; // To count channels
while (pEnumerator) {
hr = pEnumerator->Next(WBEM_INFINITE, 1, &pclsObj, &uReturn);
if (uReturn == 0) break;
VARIANT vtProp;
// Get speed (rated)
hr = pclsObj->Get(L"Speed", 0, &vtProp, 0, 0);
if (SUCCEEDED(hr) && vtProp.vt == VT_I4) {
if (static_cast<uint32_t>(vtProp.lVal) > maxSpeed) {
maxSpeed = static_cast<uint32_t>(vtProp.lVal);
}
}
VariantClear(&vtProp);
// Get configured clock speed (actual running speed)
hr = pclsObj->Get(L"ConfiguredClockSpeed", 0, &vtProp, 0, 0);
if (SUCCEEDED(hr) && vtProp.vt == VT_I4) {
if (static_cast<uint32_t>(vtProp.lVal) > maxConfiguredSpeed) {
maxConfiguredSpeed = static_cast<uint32_t>(vtProp.lVal);
}
}
VariantClear(&vtProp);
// Get capacity
hr = pclsObj->Get(L"Capacity", 0, &vtProp, 0, 0);
if (SUCCEEDED(hr) && vtProp.vt == VT_BSTR) {
totalCapacity += _wtoi64(vtProp.bstrVal);
}
VariantClear(&vtProp);
// Get memory type (SMBIOSMemoryType)
hr = pclsObj->Get(L"SMBIOSMemoryType", 0, &vtProp, 0, 0);
if (SUCCEEDED(hr) && vtProp.vt == VT_I4) {
detectedMemoryType = static_cast<uint16_t>(vtProp.lVal);
detectedType = GetDDRTypeFromSMBIOS(detectedMemoryType);
}
VariantClear(&vtProp);
// Get form factor
hr = pclsObj->Get(L"FormFactor", 0, &vtProp, 0, 0);
if (SUCCEEDED(hr) && vtProp.vt == VT_I4) {
detectedFormFactor = GetFormFactorName(static_cast<uint16_t>(vtProp.lVal));
}
VariantClear(&vtProp);
// Get data width in bits (SMBIOS "Data Width"; summed for soldered memory)
hr = pclsObj->Get(L"DataWidth", 0, &vtProp, 0, 0);
if (SUCCEEDED(hr) && vtProp.vt == VT_I4 && vtProp.lVal > 0 && vtProp.lVal <= 1024) {
dataWidthSum += static_cast<uint32_t>(vtProp.lVal);
}
VariantClear(&vtProp);
// Get bank label (for channel counting)
hr = pclsObj->Get(L"BankLabel", 0, &vtProp, 0, 0);
if (SUCCEEDED(hr) && vtProp.vt == VT_BSTR) {
uniqueBanks.insert(BstrToUtf8(vtProp.bstrVal));
}
VariantClear(&vtProp);
stickCount++;
pclsObj->Release();
}
pEnumerator->Release();
pSvc->Release();
pLoc->Release();
if (weInitializedCom) CoUninitialize();
// Fill in the info.
// WMI Win32_PhysicalMemory.Speed reports the memory data rate directly in
// MT/s (e.g. DDR4-3200 -> 3200) for all DDR generations, so no doubling is
// applied. (Previously DDR4 and earlier were doubled on the assumption WMI
// reported the I/O clock; on tested hardware it does not, which inflated
// DDR4 speeds to 2x and skewed the theoretical-bandwidth comparison.)
info.speedMT = maxSpeed;
info.configuredSpeedMT = maxConfiguredSpeed;
info.totalSticks = static_cast<uint32_t>(stickCount);
info.totalCapacityGB = totalCapacity / (1024 * 1024 * 1024);
info.type = detectedType;
info.formFactor = detectedFormFactor;
// Estimate channels from unique bank labels or stick count
// This is approximate - some systems report banks differently
if (uniqueBanks.size() >= 4u) {
info.channels = 4; // Quad channel
} else if (uniqueBanks.size() >= 2u || stickCount >= 2) {
info.channels = 2; // Dual channel (typical)
} else if (stickCount == 1) {
info.channels = 1; // Single channel
} else {
info.channels = 0; // No device data at all - unknown, not "single"
}
// Soldered memory: prefer the SMBIOS per-device data widths, which the
// device-count / locator heuristics above cannot see. Strix Halo class APUs
// report e.g. 8 x 32-bit LPDDR5X devices (256-bit bus) or a single device;
// either way the estimate is flagged so no "single-channel" warning fires
// and the iGPU comparison can correct it from the measurement.
if (IsSolderedMemory(info.type, info.formFactor)) {
info.channelsUnverified = true;
if (dataWidthSum >= 64) {
info.busWidthBits = dataWidthSum;
info.channels = std::clamp<uint32_t>(dataWidthSum / 64, 1u, 8u);
}
}
// Calculate theoretical bandwidth
// DDR bandwidth = speed (MT/s) * 8 bytes * channels / 1000 = GB/s
if (info.configuredSpeedMT > 0) {
info.theoreticalBandwidth = (info.configuredSpeedMT * 8.0 * info.channels) / 1000.0;
} else if (info.speedMT > 0) {
info.theoreticalBandwidth = (info.speedMT * 8.0 * info.channels) / 1000.0;
}
info.detected = (stickCount > 0);
return info;
}
// Format system memory info as a string for logging
std::string FormatSystemMemoryInfo(const SystemMemoryInfo& mem) {
if (!mem.detected) {
return "System Memory: Detection failed (" + mem.errorMessage + ")";
}
std::ostringstream oss;