Topic: gpu

12 stories found

Today

industry24

Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed

Perplexity shared details about its GPU-based embedding stack, including Ivy, Tulip, and ROSE, to enhance retrieval quality in AI search products by improving how cheaply embeddings can be run across indexes. This matters because optimizing embedding serving infrastructure can significantly boost the efficiency and performance of AI applications.

marktechpost.com

Yesterday

industry24

Nous Research Adds One-Click Local Model Setup to Hermes Desktop

Nou Research has simplified local model setup in Hermes Desktop to a single click, automatically optimizing and configuring the process for users based on their hardware. This update aims to make AI model deployment more accessible and efficient for end-users.

marktechpost.com

Saturday, August 29, 2026

Wednesday, August 26, 2026

releases54

Ollama releases: v0.33.1

Ollama released version 0.33.1, which includes updates to Qwen3.8 Flash Next support, cmake patches, and mlxrunner structured output for improved model loading times, highlighting ongoing development and community contributions.

github.com

Monday, August 24, 2026

newsletters48

Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye

The article discusses three key topics: the ethical considerations regarding machine rights, a new method for automating environment generation using SPADE, and improvements in GPU kernel efficiency through Hawkeye. These advancements highlight the ongoing acceleration in cybersecurity, mathematical computations, and artificial intelligence technologies.

importai.substack.com
releases48

ggml/llama.cpp releases: b10612

The ggml/llama.cpp project released version b10612, which disables a specific test for WebGPU compatibility. This update is important as it enhances the project's support for diverse hardware architectures, including Apple Silicon on macOS.

github.com

Sunday, August 23, 2026

releases48

ggml/llama.cpp releases: b10594

A code update in ggml/llama.cpp (commit b10594) optimizes device_info handling by skipping unnecessary loops when device information is not printed, reducing overhead, particularly for the CUDA backend. This improvement enhances efficiency without affecting functionality.

github.com

12 of 12 items shown. Sources: 107 days indexed.