Available at: https://digitalcommons.calpoly.edu/theses/3394
Date of Award
6-2026
Degree Name
MS in Electrical Engineering
Department/Program
Electrical Engineering
College
College of Engineering
Advisor
Andrew Danowitz
Advisor Department
Electrical Engineering
Advisor College
College of Engineering
Abstract
This thesis conducts a hardware-level analysis of the cost–benefit tradeoffs associated with increasingly complex SIMT control mechanisms in a resource-constrained GPU core. Four design implementations are evaluated: a Base Tiny GPU without warp scheduling capability; a Warp Scheduler using round-robin warp selection; a Branch Divergence implementation incorporating a dedicated divergence stack and post-dominator reconvergence mechanism; and a Dynamic Warp Allocation implementation that replaces the static warp structure with a runtime warp manager, regrouping threads by current program counter to recover SIMD lane utilization during active divergence. Each design is evaluated using two complementary measurements: functional simulation implemented using the CocoTB hardware verification framework on a matrixarithmetic test suite, and physical synthesis reports using the LibreLane RTL-toGDSII flow targeting the SkyWater 130 nm (SKY130) process node that yield cell count, die area in square micrometers, and estimated power consumption in watts. The Dynamic Warp Allocation implementation achieves the most favorable cost– benefit tradeoff among the three warp-capable designs: it is 43.6% smaller in area than the Warp Scheduler and consumes 2.8 times less power than the Branch Divergence implementation at in some configurations.