Research Engineer · AI Agents
All things agents.
I currently work at H as a researcher on Self-Evolving Computer Use Agents and more broadly how to evaluate CUAs in production settings.
I was previously at Dust, operating as a software engineer on all things agentic and building a community. Earlier, my research focus was on anomaly detection at Microsoft and how to scale multilingual NLP (back when it was cool) at Hypefactors.
Recent Posts
-
Harnesses Are Becoming State

-
Evaluation First: The Hidden Prerequisite for Self-Evolving Agents

-
LayerSkip: Early Exiting Grows up for LLMs
Why do decoder-only language models still run every token through every layer?
-
BERxiT: Early Exiting Beyond Entropy
Why should BERT trust its own confidence scores to decide when to stop thinking?
-
Instruction Fine-Tuning Evaluation and Advanced Techniques (opens in a new tab)
You fine-tuned your model to follow instructions—but how do you actually know it works? This guide unpacks the evaluation frameworks and parameter-efficient methods that separate production-ready agents from expensive science projects.