Topic: training data
3 stories found
Yesterday
COAL-SQL: Coverage-Guided Augmentation and Failure-Driven Learning for Text-to-SQL Post-Training
A new method called COAL-SQL enhances text-to-SQL translation by using coverage-guided augmentation and failure-driven learning, addressing the need for effective post-training in large language models to handle complex SQL queries. This advancement is crucial as it improves the accuracy and reliability of translating natural-language questions into executable SQL commands, essential for real-world database interactions.
Monday, September 14, 2026
Friday, September 11, 2026
Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models
The study challenges assumptions about membership evidence in language models by showing that detectable duplications are rare and hard to verify, suggesting caution when inferring training data content from prediction ease. This matters because it impacts how reliably we can use language model behavior to deduce their training data composition.
๐ฟ That's all for now. Come back tomorrow.
3 of 3 items shown. Sources: 123 days indexed.