← Back to the site

Raghav Gupta

SDET who extended test-design discipline from deterministic systems to AI agents

Profile

Five years proving out other people's systems — mobile, web, warehouse robotics — then AI agents arrived and broke the assumption test engineering was built on: that correctness is binary. At Handshake AI I design reproducible evaluation environments for coding agents; on my own time I build the agent systems themselves, the way a tester would. Looking for the role where those two halves are the same job.

What I bring

Experience

Handshake AIMar 2026 — Present · Remote
QA Engineer — AI Agent Evaluation (Contract)
  • Designed end-to-end test scenarios for AI coding agents — each packaged as a reproducible Dockerised environment with a written specification, a reference implementation, and an automated verification suite defining objective pass/fail criteria.
  • Applied boundary, negative and exploratory testing to surface agent failure modes in long-horizon planning, environment state handling, and recovery from failed tool calls.
  • Enforced determinism and eliminated flaky verification so results measure agent capability rather than environment defects.
  • Extended five years of test-design discipline to non-deterministic systems, where correctness is probabilistic and assertions must tolerate valid variation in approach while still catching wrong outcomes.
Grey OrangeSep 2025 — Dec 2025 · Gurugram, India
Software Engineer — QA / SDET
  • Owned end-to-end sprint QA — planning, test design, automation, defect tracking and retrospectives — landing on-time releases with improved defect containment.
  • Built automated regression suites for warehouse and robot workflows in PyTest, Selenium, Playwright and Page Object Model, cutting manual effort 60–70%.
  • Ran performance and load testing for high-throughput bulk operations with JMeter and REST API testing, validating SLA compliance under peak load.
  • Published pipeline metrics to Grafana for coverage, failure-trend and release-readiness visibility across stakeholders.
  • Triaged nightly automated pipelines — isolating flaky tests and coordinating fixes via JIRA — reducing false positives 25%.
ShwayJun 2023 — Sep 2025 · Remote
Software Engineer — Quality Assurance
  • Led functional, system and regression testing across iOS and Android using manual QA plus Selenium/Appium — zero critical production defects across releases.
  • Ran end-to-end API testing with Postman, RestAssured, Requests and Swagger, cutting backend defect leakage 40%.
  • Executed load testing for high-traffic flows with JMeter, improving response times under peak load 30%.
  • Designed an in-house data-driven automation framework (PyTest, Page Object Model) that reduced test maintenance 35%.
  • Integrated test execution into CI/CD with Jenkins, Git, Docker and Kubernetes for quality checks on every build.
Yellow.aiJul 2021 — Jun 2023 · Remote
Customer Success Engineer
  • Acted as quality gatekeeper for enterprise clients — owning pre-release validation, defect triage and go-live sign-offs for conversational workflows.
  • Triaged incidents, set severity, and verified fixes in staging and production, lifting CSAT and reducing repeat incidents.
  • Led technical UAT sessions with enterprise customers, converting business feedback into product improvements.

Selected projects

squared-upIndia-first expense splitter · Django/DRF + React PWAgithub.com/raghavg27/squared-up

Money core held in integer paise and verified with Hypothesis property-based tests; UPI one-tap settlement, installable PWA, Dockerised.

equity-crewMulti-agent equity research · CrewAIgithub.com/raghavg27/equity-crew

Five specialised agents fusing fundamentals, neural news search, from-scratch technical indicators and sector peer benchmarking into a validated BUY/HOLD/SELL PDF report.

git-guideAgentic RAG over GitLab docsgithub.com/raghavg27/git-guide

CrewAI + ChromaDB + Streamlit; async parallel processing and smart routing, with every answer grounded and citation-enforced against source docs.

StenoWriting standard for AI prosesteno-ai.vercel.app

A tiered editing standard that makes AI-generated prose read as though a person wrote it. Ships as a Claude skill and a portable prompt.

Toolkit

Languages & Data
Python · JavaScript · SQL · PostgreSQL · MongoDB / NoSQL · Bash
Test & Automation
Selenium · Playwright · Appium · PyTest · Hypothesis · JUnit · Page Object Model · JMeter · Postman / RestAssured
AI / Backend
CrewAI · Agentic RAG · ChromaDB · Django / DRF · React · Node.js · REST / OpenAPI
DevOps & CI/CD
Jenkins · Docker · Kubernetes · Git / GitHub Actions · Grafana · JIRA · TestRail · Confluence

Education

Degree
B.E. Computer Engineering · 2017–2021
SRM Institute of Science and Technology

Availability

Engagement
Independent contractor / B2B · EOR · full-time via an India-registered entity
Location
Delhi, India · remote worldwide, any timezone
Work authorisation
Indian citizen. No visa or sponsorship required for contract engagements.