Topic: training data

4 stories found

Yesterday

industry24

Adaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus

Adaption Labs has launched Invent a Dataset, a tool that creates training data based on task descriptions without needing a seed corpus or manual labeling, aiming to streamline the training process for machine learning models. This innovation could significantly reduce the time and effort required in preparing datasets, making model development more accessible and efficient.

marktechpost.com

Friday, September 4, 2026

research40

Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards

Benchmark contamination can inflate scores by leaking test items into training data, but its impact on reordering LLM leaderboards is limited, suggesting the reliability threat may be overstated.

arxiv.org

Tuesday, August 25, 2026

research40

Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding

Large Vision-Language Models (LVLMs), while effective in various tasks, can exhibit biased behavior due to social biases in their training data. Researchers propose a method called Counterfactual Ensemble Decoding to mitigate these biases, highlighting the importance of addressing fairness in AI systems.

arxiv.org

🌿 That's all for now. Come back tomorrow.

4 of 4 items shown. Sources: 107 days indexed.