Research

AI systems,
stress-tested.

Research-driven work exploring safety, reliability, benchmarking, and behavior under stress in large-scale software and AI systems.

Selected entries · 11

Papers & investigations.

01

AgentLens

Open-source developer tool that diagnoses AI-agent failures — classifying root causes (tool selection, loops, cascades, context pollution, state drift, overflow) and computing the wasted-token cost of each. Ships Python & TypeScript SDKs, a FastAPI server, and a CLI; integrates with Anthropic, OpenAI, LangGraph, CrewAI, AutoGen, and PydanticAI via runtime instrumentation.

Python · TypeScript · FastAPI · AI Agents

02

LLM-CODEGEN

Advanced code generation system using Large Language Models.

LLM · Generative AI · NLP

03

Large-Scale Bug Prediction

Benchmarking bug prediction across 700k Python projects.

Big Data · Research · Machine Learning

04

AI Recommender

Intelligent recommendation engine for personalized content.

Recommender Systems · Python · ML

05

AI in Software Engineering

Research on applying AI techniques to software engineering problems.

Research · AI · SE

06

Research Paper

Academic research paper repository.

Research · LaTeX

07

VLM Failure Modes

Analysis of failure modes in Vision-Language Models.

VLM · AI Safety · Research

08

VLM Adversarial Defense

Defense mechanisms against adversarial attacks on VLMs.

VLM · Adversarial ML · Python

09

Amnesic VLM Defense

Amnesic defense techniques for Vision-Language Models.

VLM · Defense · Python

10

AI Impact on Jobs

Analysis of AI's impact on the job market.

Data Analysis · Research

11

Sentiment Steering GPT

Steering GPT output sentiment.

LLM · Python · NLP

More

Research lives on GitHub.

Papers, benchmarks, and experiment code are versioned alongside the implementations.