Topic: ci
92 stories found
Yesterday
The-Art-of-Hacking/h4cker ā This repository is maintained by Omar Santos (@santosomar) and includes thousands of resources related to ethical hackin
Omar Santos maintains a repository with thousands of resources for ethical hacking and related fields like bug bounties and digital forensics, aiming to support professionals in these areas. The extensive collection covers topics such as AI security and exploit development, highlighting the importance of responsible cybersecurity practices.
What Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus
A study found that text classifiers achieve over 99% accuracy on a widely used fake news dataset, raising questions about the reliability of such high accuracy claims due to the challenges in accurately assessing fake news. The research employs a reproducible audit using simple methods to highlight potential issues with the dataset's utility for training robust models.
Tuesday, September 22, 2026
Better prompt caching for GPT-6
GPT-6 enhances prompt caching by increasing cache hit rates and introducing new diagnostics, breakpoints, and cost-reducing controls, aiming to boost efficiency and lower operational costs. This improvement is crucial as it directly impacts the performance and economic viability of using advanced language models in various applications.
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
UK AISI and EvalEval are developing methods to make benchmark results in artificial intelligence more reproducible, addressing concerns about transparency and reliability in the field. This initiative is crucial for advancing AI research by ensuring that findings can be consistently replicated and verified.

š¬ An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science
John Platt, an Oscar-winning "gigannerd" from Google, discussed using AI to automate scientific research and address climate change, emphasizing opportunities for future generations to contribute despite advancements in superintelligent technology.
Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents
The study PsyAgentBench re-runs classic psychology experiments on LLMs to assess their susceptibility to human biases without attributing those biases directly to the models, highlighting the need for contamination-aware analysis. This matters as it provides a framework to understand and mitigate potential psychological effect mimicry in AI systems.
Monday, September 21, 2026
Advisory Group on Mathematics and Artificial Intelligence
OpenAI has formed an Advisory Group on Mathematics and Artificial Intelligence to oversee the review and public communication of new AI findings. This initiative aims to ensure transparent and responsible development and dissemination of advanced AI technologies.

Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem
Researchers propose pruning large language models (LLMs) using techniques inspired by physics, specifically the Ising model for optimization problems, to improve efficiency without significantly impacting performance. This approach could lead to more resource-efficient LLMs, advancing practical applications and reducing computational costs.

AI coding has made CI a bottleneck, so we reworked ours to keep up
A company found that artificial intelligence (AI) integration had slowed down their continuous integration (CI) process, necessitating a redesign of their CI system to better accommodate the new technologies and maintain efficiency. This change is crucial as it ensures smoother software development workflows in an increasingly automated environment.
From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators
A new study proposes using large language models (LLMs) in a persona-grounded, open-ended simulation as discharge educators to better adapt to patients' literacy, recall, and personality needs, addressing limitations of current LLM evaluations that focus on static or artifact-generation tasks. This approach aims to improve patient understanding and adherence to discharge plans.
92 of 92 items shown. Sources: 124 days indexed.