Topic: quant

13 stories found

Yesterday

releases48

ggml/llama.cpp releases: b10255

The ggml/llama.cpp project updated its SYCL oneDNN SDPA to support Q4_0-Q8_0 and FP32 KV caches, extending the handling of non-FP16 key-value pairs. This update is significant as it enhances the flexibility and performance of the model in various computational scenarios.

github.com↗

Monday, August 3, 2026

research40

The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?

The study introduces the concept of "Agentic Formalism Trap" and the Evaluative Dissonance Index ($D_E$), highlighting how AI judge systems may prioritize procedural correctness over substantive truth under adversarial conditions. This research matters as it underscores potential biases in AI judicial evaluations, emphasizing the need for more nuanced approaches to ensure fairness and accuracy.

arxiv.org↗

Friday, July 31, 2026

releases48

ggml/llama.cpp releases: b10213

The ggml/llama.cpp project released version b10213, which includes support for rotated key-value cache quantization. This update enhances the efficiency and performance of the model, making it more suitable for resource-constrained environments like macOS Apple Silicon.

github.com↗

Thursday, July 30, 2026

releases2 sources⚔ Corroborated48

ggml/llama.cpp releases: b10198

The ggml/llama.cpp project released version b10198 which adds Vulkan support for quantized concat operations. This update is significant as it enhances the library's capabilities in handling quantized data with Vulkan, potentially improving performance on specific hardware platforms.

Covered by ggml/llama.cpp releases

Wednesday, July 29, 2026

releases48

ggml/llama.cpp releases: b10181

The ggml-cuda project updated its code to disable Multi-Memory Queue (MMQ) on devices with less than 48 KiB of shared memory, ensuring optimal performance across different hardware configurations. This update is crucial for maintaining compatibility and efficiency in various computational environments.

github.com↗
research40

DS@GT ARC at CheckThat! 2026: LLM-Based Trace Ranking and Grouped Reward Modeling for Multilingual Numerical Claim Verification

A new system using LLMs for trace ranking and grouped reward modeling is developed to verify multilingual numerical claims, addressing the challenge of combining language understanding with quantitative reasoning. This system is part of an entry for the CLEF 2026 CheckThat! Task 2, highlighting advancements in automated claim verification.

arxiv.org↗

Saturday, July 25, 2026

releases54

Ollama releases: v0.32.4

github.com↗

Thursday, July 23, 2026

Wednesday, July 22, 2026

🌿 That's all for now. Come back tomorrow.

13 of 13 items shown. Sources: 77 days indexed.