visit benchmark results

C++
Author

dev::author

Published

August 25, 2026

Introduction

A month ago I published a custom implementation of std::visit that stores function pointers in a flat one-dimensional array. If you haven’t read it, start there - this post builds directly on top of it.

visit implementation using flat arrays

For completeness, this is the implementation of visit using a one-dimensional array.

visit implementation using multi-dimensional arrays

A multidimensional array is an array of T elements, where the type T is a generic type - either a scalar or an array itself. Hence, a multidimensional array can be easily coded up as follows:

namespace qc{
    /**
     * @brief A light n-dimensional array.
     */
    template<typename T, std::size_t Size>
    struct poly_array{
        T m_buffer[Size] = {};

        constexpr T& operator[](std::size_t idx){
            return m_buffer[idx];
        }

        // Base case
        decltype(auto) constexpr at(std::size_t index){
            return m_buffer[index];
        }

        // Case with a pack of indices
        template<typename... Indices>
        decltype(auto) constexpr at(std::size_t index, Indices... indices){
            return m_buffer[index].at(indices...);
        }
    };
}

The .at(std::size_t, Indices...) method accepts a variadic pack of indices. As a concrete example, if I define a multi-dimensional array a of dimensions \(2 \times 3 \times 4\), a is a collection of two \(3 \times 4\) matrices. So, a.at(0,1,2) is (1,2)-th element in the 0-th matrix. Hence, we should recursively invoke a[0].at(1,2).

Next, we write a helper struct called dispatcher. The dispatcher has a static constexpr method called dispatch that invokes any callable on a pack of variants. It’s basically a smart wrapper over std::invoke, that checks the return type of the callable at compile-time and dispatches to an appropriate branch.

template<size_t... Indices>
struct dispatcher{
    template<typename Func, typename... Vs>
    static constexpr auto dispatch(Func&& f, Vs&&... vs){
        using return_type_t = decltype(std::forward<Func>(f)(std::get<0>(std::forward<Vs>(vs))...));
        if constexpr(std::is_same_v<return_type_t, void>){
            std::invoke(std::forward<Func>(f), std::get<Indices>(std::forward<Vs>(vs))...);
        }else{
            return std::invoke(std::forward<Func>(f), std::get<Indices>(std::forward<Vs>(vs))...);
        }
    };
};

Here is the complete code-listing.

visit implementation using std::mdspan

C++23 supports std::mdspan a non-owning multi-dimensional view on an underlying \(1\)D-array. I present below, a visit implementation using mdspan.

Source code repo

The source code for the above implementations can be found at github.com/quasar-chunawala/qc-visit.

Benchmarking setup

I benchmarked the qc::flat_array::visit implementation against std::visit. The code was compiled in C++ using g++ 16.2.1, with the compiler flags -std=c++23, -O2, -DNDEBUG. The benchmarks were run on Intel x86-64 Core Ultra 9 processor.

╭─ ~/repo/quantdev  main !34 ────────────────────────────────────────────────── 127 ✘  quantdev@quasar-arch  16:45:30 
╰─ lscpu
Architecture:                x86_64
  CPU op-mode(s):            32-bit, 64-bit
  Address sizes:             46 bits physical, 48 bits virtual
  Byte Order:                Little Endian
CPU(s):                      22
  On-line CPU(s) list:       0-21
Vendor ID:                   GenuineIntel
  Model name:                Intel(R) Core(TM) Ultra 9 185H
    CPU family:              6
    Model:                   170
    Thread(s) per core:      2
    Core(s) per socket:      16
    Socket(s):               1
    Stepping:                4
    Microcode version:       0x1c
    CPU(s) scaling MHz:      23%
    CPU max MHz:             5100.0000
    CPU min MHz:             400.0000
    BogoMIPS:                6144.00
    Flags:                   fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fx
                             sr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts re
                             p_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq dtes64 monitor 
                             ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc
                             _deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault ssbd ibrs ibpb 
                             stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep 
                             bmi2 erms invpcid rdseed adx smap clflushopt clwb intel_pt sha_ni xsaveopt xsavec xgetbv1 xsave
                             s split_lock_detect user_shstk avx_vnni dtherm ida arat pln pts hwp hwp_notify hwp_act_window h
                             wp_epp hwp_pkg_req hfi vnmi umip pku ospke waitpkg gfni vaes vpclmulqdq rdpid bus_lock_detect m
                             ovdiri movdir64b fsrm md_clear serialize pconfig arch_lbr ibt flush_l1d arch_capabilities
Virtualization features:     
  Virtualization:            VT-x
Caches (sum of all):         
  L1d:                       544 KiB (14 instances)
  L1i:                       896 KiB (14 instances)
  L2:                        18 MiB (9 instances)
  L3:                        24 MiB (1 instance)
NUMA:                        
  NUMA node(s):              1
  NUMA node0 CPU(s):         0-21

In order to measure performance accurately and ensure that the results of benchmarking are consistent and reproducible, I describe my setup below. Ideally when doing benchmarking, we try to disable all the potential sources of performance non-determinism in a system.

  • Disable turboboost. Intel Turbo Boost is a feature that automatically raises the CPU operating frequency when demanding tasks are running. To disable turbo in Linux, do:
# Intel
echo 1 > /sys/devices/system/cpu/intel_pstate/no_turbo
# AMD
echo 0 > /sys/devices/system/cpu/cpufreq/boost
  • Disable hyperthreading. Modern CPU cores are often made in the simultaneous multithreading(SMT) manner. It means that in one physical core you have \(2\) threads of simultaneous execution. Typically, the \(2\) threads see different architectural state(process memory, control registers state, general purpose registers state), but in reality the execution resources are shared. That means, that some other process that is scheduled on the sibling thread might steal cache space from the workload you are measuring. This can be done programatically by turning down a sibling thread in each core:
echo 0 > /sys/devices/system/cpu/cpuX/online

The pair of CPU \(N\) can be found in /sys/devices/system/cpu/cpuN/topology/thread_siblings_list.

  • Disable CPU Performance scaling and lock the frequency. CPU performance scaling enables the operating system to scale the CPU frequency up or down in order to save power or improve performance. Scaling can be done automatically with respect to the system load or be manually changed by userspace programs. The linux kernel offers CPU frequency scaling through the CPU Frequency subsystem which has \(2\) layers:
    • Scaling governors implement algorithms to compute the desired CPU frequency, potentially based off the system’s needs.
    • Scaling drivers with the CPU directly, enacting the desired frequencies that the governor is currently requesting.

A default scaling driver and governor are selected, but user-space programs such as cpupower can be used to customize this configuration.

In order to reliably benchmark the code, we do not want the CPU frequency to scale, mid-run.

We can get the governor of all the CPU cores, by running:

cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor

I set all cores to performance.

sudo cpupower frequency-set --governor performance

Since I plan to pin the benchmark run to a specific core, I will lock the frequency of that core.

sudo cpupower --cpu 3 frequency-set -d 2.3GHz -u 2.3GHz

Run cpupower -c all frequency-info | grep "current policy" -A1 to confirm that the frequency was indeed set.

  • Set scheduling policy to SCHED_FIFO and set process priority. In linux, the chrt tool is used to set or report the scheduling policy and priority.

    • -f: Use the SCHED_FIFO policy. Once a SCHED_FIFO process is running, the kernel will not preempt with another SCHED_OTHER process, it only yields the CPU when it blocks, finishes or a higher (or equal ) priority thread becomes runnable.
    • 10: The real-time priority.
  • Set CPU affinity. Processor affinity enables the binding of a process to a certain CPU core. In linux, one can do this with the taskset tool.

  • Disable address space randomization.

Address space layout randomization (ASLR) is a computer security technique involved in preventing exploitation of memory corruption vulnerabilities. In order to prevent an attacker from reliably jumping to, for example, a particular exploited function in memory, ASLR randomly arranges the address space positions of key data areas of a process, including the base of the executable and the positions of the stack, heap and libraries.

echo 0 | sudo tee /proc/sys/kernel/randomize_va_space

Benchmarking results

The benchmarking code can be found under core/tests_visit_flat_array.cpp, core/tests_visit_polyarray.cpp and core/test_visit_stdlib.cpp.

As for the benchmark itself, it consists of the following:

  • Time to visit \(N\) variant arguments. In this benchmark, a visit(visitor, args...) call is made to \(N\) variants as arguments. Each of v1,...,vN is a std::variant<Type_1, Type_2>, where all Type_is are just a simple wrapper over a double. There are no heap allocations. The floating-point values are generated randomly at run-time from a \(\text{Uniform}[0.0,1.0]\) distribution, to ensure that the compiler does not aggressively inline the result of the visit calls.

Benchmarking results