Topic: large language models

24 stories found

Today

research40

Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models

A new study evaluates large language models (LLMs) based on their coding behaviors rather than just performance metrics like pass@k, highlighting that as models improve, traditional evaluation methods become less effective in distinguishing between them.

arxiv.org

Yesterday

research40

Towards Secure Cloud-Native Computing: Unveiling Kubernetes Misconfigurations with Large Language Models

A new study using large language models highlights misconfigurations in Kubernetes, a crucial tool for cloud-native computing, emphasizing the need for enhanced security practices to protect organizational data. This research is significant as organizations increasingly rely on scalable and efficient infrastructure solutions.

arxiv.org

Friday, September 18, 2026

research40

Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations

A new method allows the discovery of stylistic variations within large language models' responses without needing supervised data or training, highlighting key stylistic dimensions based on prompts. This breakthrough could enhance understanding and control over how LLMs generate text in different styles.

arxiv.org

Thursday, September 17, 2026

research40

Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits

The study explores how large language models (LLMs) are influenced by social desirability and impression management, similar to humans during personality assessments, highlighting the need for better understanding of response distortions in AI.

arxiv.org

Wednesday, September 16, 2026

research40

Optimal Model Activation Policies for Inference Networks of Large Language Models

A new study explores optimal model activation policies for large language models to optimize performance while managing high inference costs, crucial for efficient use in NLP tasks.

arxiv.org

Friday, September 11, 2026

research40

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

The study questions the reliability of multilingual large language models (LLMs) in generating accurate text in Urdu, highlighting significant cultural and linguistic weaknesses despite these models being designed for multiple languages. This matters because it challenges the assumption that LLMs effectively support low-resource languages like Urdu, impacting their practical utility.

arxiv.org

Thursday, September 10, 2026

research40

StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

StochBench is a new Lean 4 benchmark for formal theorem proving with large language models, focusing on stochastic processes to better represent field-specific applications rather than competition math problems. This benchmark aims to improve the evaluation and application of these models in specific domains.

arxiv.org

🌿 That's all for now. Come back tomorrow.

24 of 24 items shown. Sources: 123 days indexed.