A new study evaluates large language models (LLMs) based on their coding behaviors rather than just performance metrics like pass@k, highlighting that as models improve, traditional evaluation methods become less effective in distinguishing between them.