Work / Portfolio optimizer, built twice
GPU Portfolio & Risk Decision Engine
A portfolio optimizer built twice, on CPU and GPU, with a parity suite proving both give the same answer.
- Role
- Solo build
- Year
- Jul/Aug 2026–present
- Status
- Built; CPU and GPU paths committed
- Stack
- Python, RAPIDS cuDF and cuML, NVIDIA cuOpt, CVXPY and Clarabel, pandas, Docker
- Links
- Repo
Problem
GPU speedups are easy to inflate with a slow baseline. I wanted to know where a GPU actually pays off in portfolio optimization.
What I built
A mean-variance optimizer built twice, once on CPU and once on GPU, sharing one interface. A parity suite proves both paths give the same answer before any timing is trusted, and a benchmark harness times each stage separately.
It includes three risk models, lot rounding with transaction costs, and a backtest with no lookahead.
Result
The QP solve is 2.48–2.66× faster on GPU at 3,000 assets (5.96 s to 2.41 s), with the crossover between 500 and 3,000 assets and parity to 1e-6. Lot rounding beat the greedy baseline in 18 of 18 cases. Only the solve stage wins on GPU; I don't claim an end-to-end speedup.