SSumit.
Freelance / AI Evaluation

AI Evaluation & LLM Benchmarking Specialist

Freelance AI evaluation work for Handshake AI’s Project Dynamo, focused on building software engineering benchmarks for coding-focused large language models.

AI Evaluation & LLM Benchmarking Specialist
Overview

Worked as a freelance AI evaluation specialist on Handshake AI’s Project Dynamo, designing and validating software engineering benchmarks for coding-focused large language models. The work included authoring benchmark tasks for debugging, feature implementation, repository maintenance, and code understanding; building deterministic pytest verifiers and reference implementations; calibrating benchmark difficulty against reference coding agents; and validating self-contained Docker environments through Python, Git, Linux, and CI workflows.