Open to Opportunities - Graduating May 2026


I'm
Samarth.

I build intelligent systems that bridge AI and enterprise workflows. I specialize in LLM evaluation, prompt engineering, and edge-native AI systems, focusing on production-ready pipelines, RAG architectures, and autonomous agentic workflows.

Philosophy

I engineer for measurable impact: rigorous evaluation, seamless integration, and optimized performance. The goal is reliable AI, not just impressive demos.

01

Normalized Evaluation

Metrics must be comparable. I build normalization layers to standardize disparate scores into unified scales for clear decisions.

02

Enterprise Integration

AI must fit existing workflows. I design provider-neutral pipelines that sync seamlessly with CRM and ERP systems.

03

Model Optimization

Performance matters. I use techniques like layer freezing and augmentation to prevent overfitting and ensure fast convergence.

Toolkit -
Programming & AI
PythonJavaJavaScriptC++R TensorFlowPyTorchScikit-LearnTransformersLangChainDeepEvalBERT
Development
Spring BootFastAPIFlaskReactSalesforce ApexHibernateGitDockerHeroku
Data & Cloud
SQLNumPypandasApache SparkAWS Cloudflare Workers/D1/R2ETL PipelinesSFTPStatistical AnalysisData Visualization
LLM & NLP
RAGPrompt EngineeringLLM-as-JudgeMCPGemini LlamaGPT-4NLTKSpeech RecognitionSentiment Analysis
Computer Vision
MediaPipeOpenCVComputer VisionData Augmentation
Tools & Protocols
PyQt5MIDI ProtocolD3.jsTailwind CSSPyodide/WASMV8 Isolates

Work

Selected projects that reflect my focus: clear metrics, reliable pipelines, and actionable AI insights.

PrepGenie interface screenshot
Edge AI

PrepGenie: LLM Interview Coach

Edge-native interview platform leveraging Pyodide/WASM and recursive AI failover. Delivers resume-aware RAG and real-time LLM-as-Judge scoring for contextual mock interviews.

120msAvg TTFB
<100msCold Starts
99.9%Uptime
3-ModelFailover Chain
Pyodide/WASMEdge NativepypdfLLM-as-JudgeRecursive Failover
Technical Challenge

Ensuring zero-downtime coaching sessions. I engineered a Gemini to Groq to Mistral recursive failover chain with async request handling and distributed SQL to guarantee reliable performance.

EvalGenie dashboard screenshot
Pipeline

LLM Content Evaluation Framework

Unified DeepEval pipeline benchmarking GPT-4, Gemini 2.5, and LLaMA 3-70B across 7 normalized metrics for objective, data-driven model routing.

7Metrics Normalized
96%GPT-4 Accuracy
40xGemini Cost-Efficiency
DynamicModel Routing
DeepEvalPythonBLEU/BERTScore/GEvalCustom Normalization
Production Insight

Revealed domain-specific routing strengths by benchmarking across 20+ real-world articles, enabling a production-ready pipeline that balances precision, cost, and creative output.

Penetration Testing Framework interface
Security AI

Autonomous AI Penetration Testing

Multi-agent cybersecurity framework orchestrating Reconnaissance, Vulnerability Assessment, Exploitation, and Strategic Reporting using MCP and Llama 3.3.

95%SQLi Detection
90%XSS Detection
$10K+/yrScanner Cost Saved
4 AgentsOrchestration
MCPLlama 3.3 70BCloudflare TunnelsOWASPDocker
Technical Challenge

Architected a hybrid cloud-edge pipeline that integrates industrial tools like Nmap and sqlmap via secure tunnels while enforcing strict ethical guardrails against internal IP pivoting and restricted domains.

Real-Time Stocks Dashboard
Serverless

Real-Time Stocks Serverless Dashboard

Global serverless financial dashboard utilizing V8 Isolate runtimes and React 18. Features on-demand D1 SQLite caching to prevent upstream API exhaustion.

<60sData Latency
90%API Call Reduction
ZeroIdle Compute Cost
D3.jsSVG Charting
React 18 + ViteCloudflare D1 + R2Finnhub APIOn-Demand Caching
Technical Challenge

Designed a pull-model authorization engine that intercepts requests via MAX timestamp queries. Resolves 9,999 subsequent concurrent reads entirely from edge cache within each 60-second window.

GestureSynth interface screenshot
Computer Vision

GestureSynth: Real-time Hand Gesture to MIDI

Vision-based system using Python and MediaPipe that translates hand gestures into MIDI output with low latency. Features LLM-powered adaptive tuning for optimal playability.

30 FPSPerformance
<80msLatency
40%Faster Learning
LLM-AdaptiveTuning
PythonMediaPipeOpenCVMIDIPyQt5
Technical Challenge

Calibrating gesture-to-music mapping in real-time. I built an LLM-powered adaptive tuning module that analyzes note stability metrics to automatically optimize parameters for improved playability.

Manifesto Explainer screenshot
NLP

Manifesto Explainer: Political Party Analysis

NLP platform transforming dense political manifestos into accessible insights using sentiment analysis, word clouds, and Llama-3 for contextual understanding and auto-translation.

10xFaster Scanning
ContextualUnderstanding
MultilingualAuto-Translation
DeployedWeb App
Spring BootPythonNLTKLlama-3Heroku
Technical Challenge

Manifestos are messy. I integrated Groq's Llama3-8b model to summarize what parties actually promise, turning static documents into interactive policy assistants with symbol detection.

Background

Graduating May 2026. Looking for roles where AI engineering, data science, and reliable systems actually matter.

Experience

Jun-Aug 2025

Generative AI Prompt Engineer Intern

ARMA AI Labs - California, USA

  • Established reproducible QA pipelines by building a DeepEval-powered evaluation framework leveraging 7 metrics with custom normalization for objective prompt comparison.
  • Revealed GPT-4 factual accuracy and Gemini cost-efficiency by benchmarking models across 20+ real-world articles, enabling data-driven routing strategies.
  • Delivered a production-ready model selection pipeline leveraging dynamic LLM routing to optimize content delivery across precision, cost, and creativity requirements.
DeepEvalPrompt EngineeringLLM BenchmarkingDynamic Routing
Jan 2022-Jul 2024

Senior Associate Software Developer

Genpact - India

  • Reduced manual data entry by 60% for enterprise clients by engineering scalable Salesforce CRM architectures leveraging Apex triggers and Process Builder automations.
  • Ensured low-latency data consistency across global operations by integrating Salesforce with ERP systems via SFTP and building Spring Boot microservices.
  • Resolved production defects and improved CRM data visualization by leveraging Lightning Web Components, SOQL queries, and systematic code reviews.
Salesforce CRMApexSpring BootSFTP
Jun-Aug 2021

Machine Learning Engineer Intern

TannMann Foundation - Remote

  • Achieved 96% validation accuracy in brand logo detection by optimizing a TensorFlow-based pipeline leveraging data augmentation and selective backbone freezing.
  • Mitigated overfitting in early epochs to ensure production-ready model performance under varying lighting and angle conditions.
TensorFlowComputer VisionData Augmentation

Education

Aug 2024-May 2026

MS in Data Science

Stony Brook University SUNY - New York, USA

Jun 2018-Jun 2022

BTech Computer Science & Engineering

Amity University - Noida, India

Achievements & Hackathons

GSolve.ai ML Hackathon 2023 Winner

Top Winner

Genpact Internal Hackathon

Dare in Reality DS Hackathon

Rank 159

Data Science Competition

Purchase Pattern Prediction

Rank 123 / 2000+

Univ.ai

Contact

Reach out if you're hiring for AI Engineering, Data Science, or Backend Development roles.