Topic: llama

59 stories found

Yesterday

releases2 sources⚡ Corroborated48

ggml/llama.cpp releases: b10214

ggml/llama.cpp released version b10214, which includes an update to add n_embd_head. This update is significant for users needing enhanced embedding dimensions in their models, particularly on Apple Silicon Macs.

Covered by ggml/llama.cpp releases
releases48

ggml/llama.cpp releases: b10216

The ggml/llama.cpp project updated its Vulkan backend by adding support for 1D pooling operations, including necessary data structures and a compute shader, to enhance computational efficiency in certain machine learning tasks.

github.com

Thursday, July 30, 2026

releases4 sources🛡️ Verified48

ggml/llama.cpp releases: b10199

The ggml/llama.cpp project released version b10199, which adds support for input embedding to generate the next token and fixes issues with server_batch(). These updates enhance the model's functionality and performance.

Covered by ggml/llama.cpp releases
releases48

ggml/llama.cpp releases: b10197

The ggml/llama.cpp project added support for alternative convolution layouts, enhancing flexibility and ensuring compatibility across different computational kernels. This update is crucial for improving the library's versatility in handling various neural network architectures.

github.com

Wednesday, July 29, 2026

releases2 sources⚡ Corroborated48

ggml/llama.cpp releases: b10182

The ggml/llama.cpp project released a new version addressing security issues by moving suppress_tokens handling to common/sampling and removing has_logit_bias, emphasizing improved safety in the latest update.

Covered by ggml/llama.cpp releases
releases48

ggml/llama.cpp releases: b10181

The ggml-cuda project updated its code to disable Multi-Memory Queue (MMQ) on devices with less than 48 KiB of shared memory, ensuring optimal performance across different hardware configurations. This update is crucial for maintaining compatibility and efficiency in various computational environments.

github.com
research40

TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking

TimeCapsule is a new method using a large language model to address the temporal bias in contemporary training data, making these models unreliable for historical contexts. The researchers developed TimeCapsule, a 1.2B-parameter model, to better handle historical sensemaking by reducing present-day concept encoding.

arxiv.org

Tuesday, July 28, 2026

releases2 sources⚡ Corroborated48

ggml/llama.cpp releases: b10173

The ggml/llama.cpp project has released version b10173, which includes the addition of the Laguna-S-2.1 language model. This update is significant for developers and users interested in advanced text generation capabilities.

Covered by ggml/llama.cpp releases

Monday, July 27, 2026

releases54

Ollama releases: v0.32.5

github.com

59 of 59 items shown. Sources: 73 days indexed.