Research Engineer · AI Agents
All things agents.
I work at H on Computer Use Agents and especially how to evaluate them in production settings.
I previously worked at Dust, where I built product infrastructure and community. Earlier, I worked on anomaly detection at Microsoft and multilingual NLP at Hypefactors, including systems serving about a billion inferences a day.
Recent Posts
-
Evaluation First: The Hidden Prerequisite for Self-Evolving Agents

-
LayerSkip: Early Exiting Grows up for LLMs
Why do decoder-only language models still run every token through every layer?
-
BERxiT: Early Exiting Beyond Entropy
Why should BERT trust its own confidence scores to decide when to stop thinking?
-
Instruction Fine-Tuning Evaluation and Advanced Techniques (opens in a new tab)
You fine-tuned your model to follow instructions—but how do you actually know it works? This guide unpacks the evaluation frameworks and parameter-efficient methods that separate production-ready agents from expensive science projects.
-
DeeBERT: Teaching BERT When to Stop Thinking
Why does BERT need twelve layers to classify “I love this movie” as positive?