StarCoder came out of BigCode, a collaboration between Hugging Face and ServiceNow Research built around training genuinely open, auditable code models rather than closed ones trained on undisclosed data. The original 15.5 billion parameter model used multi-query attention and fill-in-the-middle training on a trillion tokens from The Stack, a curated dataset spanning more than 80 programming languages, and a Python fine-tune of it reached around 40% pass@1 on HumanEval, competitive with much larger commercial models at the time.
StarCoder2 followed with models at 3B, 7B, and 15B parameters, trained on three to four trillion tokens across more than 600 programming languages, using grouped query attention and an extended 16K token context window. The flagship 15B model reaches or beats the capability of larger proprietary models despite its comparatively small size, which has made it a popular choice for teams wanting a strong open code model without needing to run something enormous.
Being fully open access, including training data provenance through The Stack, has made StarCoder a reference point for discussions about transparency in code model training, not just a benchmark-chasing release.