A new study finds that AI benchmarks are plateauing, indicating a need for more diverse evaluation methods to drive continued progress in artificial intelligence research. This matters because it highlights potential limitations in current benchmarking practices and suggests the field must evolve to foster innovation.