SSumit.
Experience

Where I've been.

Work
  1. Jul 2026Present

    Al Evaluation & LLM Benchmarking Specialist.

    Handshake AI (Freelance, Remote) · July 2026 – Present

    Worked as a Freelance AI Evaluation Engineer on Handshake AI's Project Dynamo, designing software engineering benchmarks for coding-focused large language models. Built deterministic evaluation pipelines, calibrated benchmark difficulty, and validated Docker-based benchmark environments to support reliable model evaluation.

    • Authored software engineering benchmark tasks covering debugging, feature implementation, repository maintenance, and code understanding.
    • Developed deterministic pytest verifiers and reference implementations for automated evaluation.
    • Calibrated benchmark difficulty against reference coding agents to ensure rigorous and reproducible evaluation.
    • Built and validated self-contained Docker environments for reproducible benchmark execution.
    • Performed end-to-end validation using Python, Git, Linux, Docker, and CI workflows before submission.
    • Evaluated and benchmarked coding-focused large language models to improve reasoning, debugging, and code generation capabilities.
    PythonGitGitHubLinuxDockerPytestCI/CDBashJSONYAMLMarkdownSoftware DebuggingBenchmark EngineeringAI EvaluationLarge Language Models (LLMs)Terminal-Based Development
  2. Aug 2024Sep 2024

    Data Science and Machine Learning Intern

    YBI Foundation · Remote

    Completed a Data Science and Machine Learning internship, building a Movie Recommendation System end-to-end.

    • Built a hybrid recommender using collaborative and content-based filtering.
    • Preprocessed data and engineered features for movie metadata.
    • Explored NLP techniques for computing similarity across titles and descriptions.
    • Documented methodology, evaluation metrics and results.
    PythonPandasNumPyScikit-learnNLP