Topic: training data

3 stories found

Yesterday

research40

COAL-SQL: Coverage-Guided Augmentation and Failure-Driven Learning for Text-to-SQL Post-Training

A new method called COAL-SQL enhances text-to-SQL translation by using coverage-guided augmentation and failure-driven learning, addressing the need for effective post-training in large language models to handle complex SQL queries. This advancement is crucial as it improves the accuracy and reliability of translating natural-language questions into executable SQL commands, essential for real-world database interactions.

arxiv.orgโ†—

Friday, September 11, 2026

research40

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

The study challenges assumptions about membership evidence in language models by showing that detectable duplications are rare and hard to verify, suggesting caution when inferring training data content from prediction ease. This matters because it impacts how reliably we can use language model behavior to deduce their training data composition.

arxiv.orgโ†—

๐ŸŒฟ That's all for now. Come back tomorrow.

3 of 3 items shown. Sources: 123 days indexed.