Normalized Evaluation
Metrics must be comparable. I build normalization layers to standardize disparate scores into unified scales for clear decisions.
I build intelligent systems that bridge AI and enterprise workflows. I specialize in LLM evaluation, prompt engineering, and edge-native AI systems, focusing on production-ready pipelines, RAG architectures, and autonomous agentic workflows.
I engineer for measurable impact: rigorous evaluation, seamless integration, and optimized performance. The goal is reliable AI, not just impressive demos.
Metrics must be comparable. I build normalization layers to standardize disparate scores into unified scales for clear decisions.
AI must fit existing workflows. I design provider-neutral pipelines that sync seamlessly with CRM and ERP systems.
Performance matters. I use techniques like layer freezing and augmentation to prevent overfitting and ensure fast convergence.
Selected projects that reflect my focus: clear metrics, reliable pipelines, and actionable AI insights.
Edge-native interview platform leveraging Pyodide/WASM and recursive AI failover. Delivers resume-aware RAG and real-time LLM-as-Judge scoring for contextual mock interviews.
Ensuring zero-downtime coaching sessions. I engineered a Gemini to Groq to Mistral recursive failover chain with async request handling and distributed SQL to guarantee reliable performance.
Unified DeepEval pipeline benchmarking GPT-4, Gemini 2.5, and LLaMA 3-70B across 7 normalized metrics for objective, data-driven model routing.
Revealed domain-specific routing strengths by benchmarking across 20+ real-world articles, enabling a production-ready pipeline that balances precision, cost, and creative output.
Multi-agent cybersecurity framework orchestrating Reconnaissance, Vulnerability Assessment, Exploitation, and Strategic Reporting using MCP and Llama 3.3.
Architected a hybrid cloud-edge pipeline that integrates industrial tools like Nmap and sqlmap via secure tunnels while enforcing strict ethical guardrails against internal IP pivoting and restricted domains.
Global serverless financial dashboard utilizing V8 Isolate runtimes and React 18. Features on-demand D1 SQLite caching to prevent upstream API exhaustion.
Designed a pull-model authorization engine that intercepts requests via MAX timestamp queries. Resolves 9,999 subsequent concurrent reads entirely from edge cache within each 60-second window.
Vision-based system using Python and MediaPipe that translates hand gestures into MIDI output with low latency. Features LLM-powered adaptive tuning for optimal playability.
Calibrating gesture-to-music mapping in real-time. I built an LLM-powered adaptive tuning module that analyzes note stability metrics to automatically optimize parameters for improved playability.
NLP platform transforming dense political manifestos into accessible insights using sentiment analysis, word clouds, and Llama-3 for contextual understanding and auto-translation.
Manifestos are messy. I integrated Groq's Llama3-8b model to summarize what parties actually promise, turning static documents into interactive policy assistants with symbol detection.
Graduating May 2026. Looking for roles where AI engineering, data science, and reliable systems actually matter.
ARMA AI Labs - California, USA
Genpact - India
TannMann Foundation - Remote
Stony Brook University SUNY - New York, USA
Amity University - Noida, India
Top Winner
Genpact Internal Hackathon
Rank 159
Data Science Competition
Rank 123 / 2000+
Univ.ai
Reach out if you're hiring for AI Engineering, Data Science, or Backend Development roles.