A new benchmark called BBOWP-Bench evaluates large language models' ability to solve black-box optimization word problems, highlighting the need for better automatic formulation techniques in optimization due to the critical impact of problem formulation on solution quality.