We Benchmarked AI Models on Git Tasks. Results Surprised Us

Most AI model benchmarks measure general coding ability or reasoning. GitBench, built by GitKraken developer advocate Chris Griffing, measures something narrower and more practical: how well a given AI model handles specific Git tasks, starting with commit squashing, identifying which commits in a messy history should be combined into one clean commit.

In this clip, Chris demos GitBench live and walks through a few findings that don't show up anywhere else. One model handled commit squashing perfectly with structured JSON output, but performed noticeably worse with plain text output on the exact same task. Another got steadily worse, not better, as its reasoning effort setting was increased, essentially overthinking itself into a wrong answer.

The bigger point: general-purpose benchmarks can't catch this kind of detail, because models can end up optimizing for the benchmarks everyone already measures. A narrower, task-specific benchmark like GitBench surfaces the kind of practical, unintuitive result that actually helps you pick the right model and settings for a real Git task, instead of guessing.

GitBench is still early and actively evolving, with more benchmarks and models being added regularly.

Check out GitBench at gitbench.gitkraken.com.

GitKraken Desktop:
http://tr.ee/GKDYT

GitKraken CLI:
http://tr.ee/CLIYT

GitLens for VS Code:
http://tr.ee/GLYT

Git Integration for Jira:
http://tr.ee/GijYT

Git Blog:
http://gitkraken.com/blog