Topic: fast
6 stories found
Today
Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed
Perplexity shared details about its GPU-based embedding stack, including Ivy, Tulip, and ROSE, to enhance retrieval quality in AI search products by improving how cheaply embeddings can be run across indexes. This matters because optimizing embedding serving infrastructure can significantly boost the efficiency and performance of AI applications.
Friday, September 4, 2026
Show HN: TERMy – A fast terminal assistant that does not use LLMs
TERMy is a new terminal assistant tool designed to enhance productivity without relying on large language models, offering faster and more direct command assistance. Its significance lies in providing an alternative for users seeking efficient terminal interactions free from the potential biases and limitations associated with LLMs.
Tuesday, September 1, 2026
Wednesday, August 26, 2026
How loveholidays is making everyone a builder with Codex
loveholidays employs OpenAI Codex to democratize software development, enabling broader team participation and accelerating product creation. This approach highlights how advanced AI tools can make complex technical tasks more accessible, enhancing efficiency and innovation within businesses.
Tuesday, August 25, 2026
Jalapeño’s first results show industry-leading speed and efficiency in AI inference
Jalapeño, a new custom inference chip from OpenAI, offers significantly faster and more energy-efficient AI processing. This breakthrough could revolutionize the industry by enhancing model throughput and reducing latency in modern applications.
Monday, August 24, 2026
Self-Speculation for Faster Reasoning Models
A new approach called "self-speculation" has been proposed to enable large language models to generate faster reasoning processes without sacrificing quality, addressing the need for quicker decision-making in complex tasks. This development is crucial as LLMs are increasingly used in scenarios requiring rapid and accurate multi-step reasoning.
🌿 That's all for now. Come back tomorrow.
6 of 6 items shown. Sources: 107 days indexed.