Performance Optimizations
Tracked optimizations applied to the library. Each entry records what was changed, why, the measured impact, and any portability considerations.
The baseline for all measurements is the state before any optimization was applied (pipelines, shader modules, descriptor set layouts, and pipeline layouts created and destroyed on every operation call).
Optimization 1: Filter Pipeline Caching
Date: 2026-07-26
Status: Applied
What changed: Before this change, vc_filter created and destroyed
all Vulkan pipeline state on every invocation: three shader modules
(vkCreateShaderModule / vkDestroyShaderModule), three compute
pipelines (vkCreateComputePipelines / vkDestroyPipeline), two
descriptor set layouts, two pipeline layouts. This is a driver-level
compilation step with non-trivial CPU cost.
After: pipeline state is created once on the first vc_filter call and
cached in a static VcFilterPipelineCache struct. Subsequent calls reuse
the cached state.
Vulkan primitives affected:
vkCreateShaderModule— 3 calls per invocation → 3 calls totalvkCreateComputePipelines— 3 calls per invocation → 3 calls totalvkCreateDescriptorSetLayout— 2 calls per invocation → 2 calls totalvkCreatePipelineLayout— 2 calls per invocation → 2 calls total
Portability: Fully portable. Uses standard Vulkan 1.3 API. No vendor extensions required.
Optimization 2: GPU-Only Buffer Allocation (VMA Flags)
Date: 2026-07-26
Status: Applied
What changed: prefixes_buf and output_buf in vc_filter
switched from VMA_ALLOCATION_CREATE_HOST_ACCESS_RANDOM_BIT to
VMA_MEMORY_USAGE_AUTO (flags=0). Neither buffer is accessed by
the CPU. Staging buffers (totals_staging, offsets_staging) retain
host access for the pass-2 CPU merge.
Why it matters: BAR memory (HOST_ACCESS_RANDOM) is limited to
~256 MB on discrete GPUs and has lower bandwidth than device-local VRAM.
Intermediate GPU-only buffers that exceed BAR cause the driver to spill
into system RAM — a sudden performance cliff (GPU reads/writes at PCIe
speed instead of HBM). Device-local allocation prevents this entirely.
Performance impact: Not a flat percentage — a scalability guard. Below ~64 MB columns there is no difference. Above that, avoids a ≥10x slowdown when BAR is exhausted.
Portability: Fully portable. VMA_MEMORY_USAGE_AUTO lets VMA
pick the best memory type for each GPU.
Optimization 3: Reusable Command Buffer per Context
Date: 2026-07-26
Status: Applied
What changed: vc_filter previously allocated a command buffer
via vkAllocateCommandBuffers and freed it via
vkFreeCommandBuffers on every invocation. Now a single
VkCommandBuffer is allocated once in vc_create_context (as
VcContext::reusable_cmd), freed in vc_destroy_context, and
reused across all operations via vkResetCommandBuffer.
Why it matters: Command buffer allocation hits the driver allocator
on every call. The pattern also establishes a convention for all
future operations — they should use ctx->reusable_cmd instead of
allocating their own.
Performance impact: Not a measurable speedup for large columns (kernel time dominates), but eliminates two driver allocator calls per invocation and makes repeated small-column filter calls faster.
Vulkan primitives affected:
vkAllocateCommandBuffers— 1 call pervc_filterinvocation → 0vkFreeCommandBuffers— 1 call pervc_filterinvocation → 0
Portability: Fully portable. Uses standard vkResetCommandBuffer
with VK_COMMAND_POOL_CREATE_RESET_COMMAND_BUFFER_BIT (already set
on the context command pool). Vulkan 1.0.
Optimization 4: Subgroup Scan in Filter Pass 1
Date: 2026-07-26
Status: Applied
What changed: filter_pass1.comp used a full shared-memory Blelloch
scan (upsweep + downsweep, 8 barriers per workgroup). On GPUs with
VK_KHR_shader_subgroup_arithmetic, a variant shader
(filter_pass1_subgroup.comp) now uses subgroupExclusiveAdd()
for intra-wave scan — register-speed, 2 barriers — with shared memory
only for the 4-way cross-wave merge. Selection at context creation
via ctx->has_subgroup_arithmetic; falls back to Blelloch otherwise.
Vulkan primitives affected:
VK_KHR_shader_subgroup_arithmeticdevice extensionVkPhysicalDeviceSubgroupProperties::subgroupSizequeried atvc_create_context
Estimated performance: 20-30% faster scan phase (8 barriers → 2). Benchmarks pending — scan time is a fraction of end-to-end filter (dominated by PCIe transfers).
Portability: VK_KHR_shader_subgroup_arithmetic is supported on
AMD (Vulkan 1.1+), NVIDIA (Maxwell+, Vulkan 1.1+), Intel (Gen12+),
and MoltenVK (Metal 3.0+). Fallback to Blelloch for others.
Optimization 5: Push Descriptors in vc_filter
Date: 2026-07-26
Status: Applied
What changed: vc_filter allocated a descriptor pool and three
descriptor sets on every call, then called vkUpdateDescriptorSets +
vkCmdBindDescriptorSets for each pass. With
VK_KHR_push_descriptor, descriptors are written directly into the
command buffer via vkCmdPushDescriptorSet — no pool, no sets, no
allocate/update/bind/destroy calls.
Vulkan primitives affected:
vkCreateDescriptorPool— 1 call per invocation -> 0vkAllocateDescriptorSets— 3 -> 0vkUpdateDescriptorSets— 3 -> 0vkCmdBindDescriptorSets— 3 -> 0vkDestroyDescriptorPool— 1 -> 0vkCmdPushDescriptorSet— 3 new calls (push path only)VK_KHR_push_descriptordevice extension
Performance impact: Eliminates 11 driver calls per filter invocation. Not measurable for large columns (kernel time dominates) but reduces CPU overhead on repeated or small-column calls.
Portability: Supported on AMD, NVIDIA, Intel, and MoltenVK (Vulkan
1.1+). Devices without the extension fall back to the traditional
pool+set path automatically. Feature detection at vc_create_context
stores ctx->has_push_descriptors.
Optimization 6: Sort Pipeline Caching
Date: 2026-07-26
Status: Applied
What changed: vc_sort creates pipeline state (3 shader modules, 3
compute pipelines, 3 descriptor set layouts, 3 pipeline layouts) once on the
first call and caches them in ctx->sort_cache. This is the same
context-scoped pattern used by vc_filter and vc_reduce.
Vulkan primitives affected:
vkCreateShaderModule— 3 calls pervc_sortinvocation → 3 calls totalvkCreateComputePipelines— 3 per invocation → 3 totalvkCreateDescriptorSetLayout— 3 per invocation → 3 totalvkCreatePipelineLayout— 3 per invocation → 3 total
Portability: Fully portable. Standard Vulkan 1.3. No vendor extensions.
Planned Optimizations
Subgroup-size-aware workgroup tuning. Query
VkPhysicalDeviceSubgroupProperties::subgroupSizeat context creation. Use to enforce workgroup sizes that are multiples of the native wave/warp size (AMD=64, NVIDIA=32, Intel=32). Current 256 threads works for all vendors but may not be optimal for occupancy.