Topic: framework

16 stories found

Friday, September 4, 2026

research40

Counterexamples as Feedback for Agent Self-Correction

A new framework called A-CEGIS has been developed to help artificial intelligence agents self-correct by using counterexamples as feedback, addressing limitations in current single-turn metrics which fail to assess the agents' ability to repair mistakes. This matters because it enhances the reliability and adaptability of AI systems in real-world applications where initial errors need correction.

arxiv.org

Thursday, September 3, 2026

ai_labs75

Safety overview: GPT-6 Astra

GPT-6 Astra has achieved the Critical level of cybersecurity, making it the company's most secure widely used model. This milestone underscores the firm's commitment to enhancing safety in their advanced language models.

openai.com

Wednesday, September 2, 2026

Friday, August 28, 2026

research40

Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework

A new framework using interpretable large language models aims to automate pricing decisions in the tourism industry by converting complex, unstructured travel orders into executable policies, addressing the limitations of traditional rule engines. This advancement is crucial as it can lead to more efficient and adaptable pricing strategies in a rapidly changing market.

arxiv.org

Thursday, August 27, 2026

trending42

AI Engineer Notebooks – free, framework-free RAG/agents/evals on Colab

AI engineers now have access to free, framework-independent resources for building and testing Retrieval-Augmented Generative (RAG) agents and evaluators directly in Google Colab notebooks. This development is significant as it lowers barriers to entry for experimenting with advanced AI models without needing extensive setup or specific tooling knowledge.

github.com

Wednesday, August 26, 2026

research35

FLARE: A Systematic, Uncertainty-Aware Framework for Evidence-Based Adoption of Artificial Intelligence in Healthcare

A new framework called FLARE aims to assess the economic viability of AI in real clinical settings, addressing the gap in current evaluations that focus solely on model accuracy. This systematic approach is crucial for ensuring that AI adoption in healthcare is both effective and cost-beneficial.

arxiv.org

Monday, August 24, 2026

research40

When Do LLMs Replace Fine-Tuned NLU? A Decision Framework for Intent Detection in Production Conversational Systems

The study compares zero-shot large language models (LLMs) to fine-tuned natural language understanding (NLU) classifiers for intent detection, concluding that the suitability of LLMs varies depending on the specific intent space. On comprehensive datasets like ATIS and CLINC150, the performance of LLMs relative to fine-tuned NLU classifiers depends on the context and type of intents involved.

arxiv.org

16 of 16 items shown. Sources: 107 days indexed.