About
I work at the compiler and runtime layer of ML systems: the part that decides whether a model runs at the speed the hardware allows. I am interested in how intermediate representations, memory planning, and performance models can drive code generation for heterogeneous hardware.
For my M.S. thesis at Carnegie Mellon (ECE), I am working on an ILP-based tensor memory planner for state-space models, which extends ILP memory allocation with cyclic lifetimes so that recurrent state is planned rather than kept live everywhere. Before that I built TransISA, a static x86-to-AArch64 binary translator on the LLVM C++ API, published at the 2026 Improving Scientific Software Conference. I am a core developer and Community Council member of sktime, a maintainer of pytorch-forecasting, and a contributor to Apache TVM. Before CMU I completed a B.Tech in Computer Engineering at Pandit Deendayal Energy University and worked as an ML engineer at Unify on the Ivy framework.
In my free time I play football and share memes with friends.
Publications
-
TransISA: A Static Transpiler for Migrating Legacy x86 Assembly to ARM via LLVM IR
Felix Hirwa Nshuti, Shakti Mishra
Proceedings of the 2026 Improving Scientific Software Conference, pp. 1–6. NCAR/UCAR (NCAR/TN-593+PROC), 2026.
Presentations
-
TransISA: A Static Assembly Transpiler for Automating x86-to-ARM Migration in Scientific Computing
Improving Scientific Software Conference, Boulder, CO — 2026
Selected Projects
-
Enhancing Prophet with Question-Aware Captioning for Knowledge-Based VQA
Python, PyTorch, Transformers — Aug 2025 to Dec 2025
- Improved knowledge-based VQA performance by integrating question-aware captioning.
- Built a module generating contextually relevant captions to improve knowledge retrieval.
-
Context-Aware Demand Forecasting in Pittsburgh's Bike Share System
Python, XGBoost, CatBoost, statsmodels — Aug 2025 to Dec 2025
- Incorporated temporal, weather, and event-based features to improve prediction accuracy.
- Applied time series analysis to capture complex usage patterns.
-
TransISA — Lightweight CISC-to-RISC Transpiler
LLVM, x86, AArch64, C++ — Dec 2024 to May 2025
- Built a compiler pipeline using the LLVM C++ API to translate x86 to AArch64 via LLVM IR.
- Extracted and analyzed control-flow graphs to support instruction-level translation.
-
Scaling Deep Learning Backends with sktime
Python, PyTorch, TensorFlow, Darts — May 2024 to Aug 2024
- Added GRU and GRU-FCNN classifiers to sktime using PyTorch.
- Migrated classifier models from legacy sktime-dl into the main sktime repository.
- Implemented the modular interface for darts regression models in sktime.