Datasets:

Modalities:
Tabular
Text
Formats:
parquet
Languages:
code
ArXiv:
Tags:
code
License:

Benchmark results on downstream coding models?

#20
by Recandle - opened

Has anyone evaluated how much stack-v3-train improves downstream coding models compared to stack-v2? I'm especially interested in HumanEval, LiveCodeBench, or SWE-bench results.

Sign up or log in to comment