RSSAmplifier

Blog

Ash's Blog

Recent content on Ash's Blog

ashvardanian.comRSS feed ↗48 posts

Latest posts

NumKong: 2'000 Mixed Precision Kernels For All 🦍

Around 2'000 SIMD kernels for mixed-precision BLAS-like numerics — dot products, batched GEMMs, distances, geospatial, ColBERT MaxSim, and mesh alignment — from Float6 to Float118, leveraging RISC-V, Intel AMX, Arm SME, and WebAssembly Relaxed SIMD, in 7 languages and 5 MB.

Full Unicode Search at 50× ICU Speed with AVX‑512

ICU gets Unicode right and pays for it. This post shows a different approach: fold-safe windows, SIMD probes, and verifiers for fast UTF‑8 search.

Tuning TLS: AES-256 Beats ChaCha20 on Every CPU

AES-256-GCM now beats ChaCha20-Poly1305 by up to 3x on every modern CPU with hardware acceleration, reversing the 2015 mobile performance advice.

Scaling Elections with GPUs and Mojo across Nvidia and AMD 🔥

Taking the computationally expensive Schulze voting method from theory to practice on GPUs — exploring parallel algorithms, hardware optimizations, and why Mojo surprised me.

2x Faster Hashes on AWS Graviton: NEON → SVE2

SVE2 AES on AWS Graviton 4 delivers 2x faster string hashing than NEON, but SVE's variable-length promise fell flat: Graviton 4 regressed to 128-bit vectors.

Before AI's Kepler Moment - Are LLMs the Epicycles of Intelligence?

Like Ptolemaic astronomers stacking circles to predict planetary motion, we stack transformer layers to approximate intelligence—accurate yet awaiting our Kepler moment.

How a String Library Beat OpenCV at Image Processing by 4x

StringZilla's SIMD-optimized Look-Up Tables beat OpenCV by 4x on image processing, proving string manipulation techniques excel beyond text.

Processing Strings 109x Faster than Nvidia on H100

StringZilla v4 brings CUDA acceleration for string processing: 109x faster than Nvidia's CuDF on edit distances, plus 52-bit MinHash fingerprinting and AES-based hashing.

Beyond OpenMP in C++ & Rust: Taskflow, Rayon, Fork Union 🍴

Most C and Rust thread pools run 10x slower than OpenMP on fork-join workloads. Fork Union closes the gap to 20% with 300 lines, dodging mutexes and CAS.

CUDA Hello World: Done Less Wrong

Moving beyond triple-bracket kernel launches to production-ready CUDA: proper error handling, CUDA Driver API, cooperative groups, and inline PTX for robust GPU code.

The Longest Nvidia PTX Instruction

Exploring the 79-character mma.sp::ordered_metadata instruction for Tensor Cores on Hopper H100s, plus insights on PTX, SASS, and the evolving complexity of GPU ISAs.

Hiding x86 Port Latency for 330 GB/s/core Reductions 🫣

By exploiting CPU port differences between Intel and AMD, interleaving FMA with addition instructions boosts AVX-512 reduction kernels from 211 GB/s to 330 GB/s per core.

Parsing JSON in C & C++: Singleton Tax

Custom allocators can dramatically reduce JSON parsing overhead, but singleton dependencies lurk everywhere—from malloc to std::isspace—limiting multi-threaded gains.

10x Faster C++ String Split, 16 Years Later 👴🏻

A 16-year-old StackOverflow question on string splitting gets a modern answer: SIMD-accelerated tokenization that's 10x faster than STL and cleaner than ranges.

The Next 31 Years of Developing Unum

Nine years into a 40-year commitment to AI infrastructure. Reflecting on mistakes, open-source wins, and why the real work is just beginning.

Understanding SIMD: Infinite Complexity of Trivial Problems 🔥

Why SIMD programming is harder than it looks: exploring cosine similarity optimization across x86 and Arm architectures, from basic AVX2 to cutting-edge SVE2 instructions.

5x Faster Set Intersections: SVE2, AVX-512, & NEON 🤐

Achieving 5x faster set intersections on AWS Graviton 4 using Arm's SVE2 and Intel's AVX-512 with specialized instructions like HISTCNT, MATCH, and VP2INTERSECT.

35% Discount on Keyword Arguments in Python 🐍

Keyword arguments in Python aren't free. Switching from PyArg_ParseTupleAndKeywords to METH_FASTCALL and manual parsing yields a 35% speedup in function calls.

NumPy vs BLAS: Losing 90% of Throughput

NumPy's binding overhead wastes up to 90% of BLAS throughput on dot products; SimSIMD recovers that performance with optimized Python interfaces.

The Painful Pitfalls of C++ STL Strings 🧵

C STL strings bring 20,000 lines of code to every compilation, yet remain slower than LibC and error-prone. StringZilla offers a faster, more intuitive alternative.

USearch Molecules: 28 Billion Chemical Embeddings on AWS ⚗️

Indexing 28 billion chemical embeddings from 7 billion molecules using optimized Jaccard distance kernels, achieving 99% recall at 3,700 QPS on AWS Open Data.

Binding a C++ Library to 10 Programming Languages 🔟

Binding USearch to 10 languages reveals the friction in each ecosystem—from Python's simplicity to Java's complexity—and why C still rules performance-critical code.

Python, C, Assembly - 2'500x Faster Cosine Similarity 📐

From pure Python to AVX-512 assembly, optimizing cosine similarity reveals a 2,500x speedup through SIMD, FP16, and VNNI instructions on modern Intel CPUs.

GCC Compiler vs Human - 119x Faster Assembly 💻🆚🧑‍💻

Hand-written AVX-512FP16 SIMD code achieves 119x speedup over GCC's auto-vectorization for Jensen-Shannon divergence, using reciprocal square roots and FMA instructions.

Accelerating JavaScript arrays by 10x for Vector Search 🏹

Supercharging JavaScript's TypedArray with C bindings and SIMD achieves 10x speedup over native Arrays for AI vector operations in Node.js.

Our CPython bindings got 5x faster without PyBind11 🐍

Switching from PyBind11 to direct CPython C API bindings reduced StringZilla's call latency by 5x, proving that understanding CPython internals pays dividends.

SciPy distances... up to 200x faster with AVX-512 & SVE 📏

AVX-512 and Arm SVE bring masked loads and native f16 math to vector distances. SimSIMD exploits both, running up to 200x faster than SciPy on modern CPUs.

Combinatorial Stable Marriages for DBMS Semantic Joins 💍

Replacing billion-entry preference lists with vector search indexes makes the Nobel Prize-winning Stable Marriage algorithm scale to real-world dating and database joins.

StringZilla: 5x faster strings with SIMD & SWAR 🦖

StringZilla uses SIMD tricks to hit 16 GB/s substring search, beating standard libraries by 5-10x and parsing multi-terabyte files Python couldn't handle.

Abusing Vector Search for Texts, Maps, and Chess ♟️

Vector search isn't just for AI embeddings. Use HNSW for geo-spatial queries, stock covariance, chess positions, text tokens, and multi-modal recommendations with custom metrics.

Counting Strings in C++: 30x Throughput Difference 💬

From junior to expert: four C solutions to count unique strings reveal 30x performance gaps. Proficiency dramatically impacts even trivial single-threaded code.

We went through life with a smile 💔

A love story cut short — remembering Sona, who taught me what it means to truly care for others, build with purpose, and find someone who shares your frequency.

Mastering C++ with Google Benchmark ⏱️

Write faster C with Google Benchmark: measure everything from math micro-ops to sorting algorithms, and discover 60x speedups with simple compiler tricks.

Failing to Reach DDR4 Bandwidth 🚌

Why reaching DDR4's theoretical bandwidth is nearly impossible. CPU parallel reductions struggle to hit 60% saturation while GPU solutions easily exceed 79%.

Crushing CPUs with 879 GB/s Reductions in CUDA

GPU code beats optimized CPU parallel reductions by 10x, reaching 879 GB/s. CUB achieves 94% bandwidth saturation while CPU barely hits 60%.

Apple to Apple Comparison: M1 Max vs Intel 🍏

M1 Max with DDR5 challenges 64-core server workloads in hash-table benchmarks. Memory bandwidth matters more than core count for many real-world tasks.

Hyperscaler Shopping List: 2022 Data Center Tech Frenzy ☁️

DDR5, PCIe Gen5, CXL, and new CPUs from Intel, AMD, and NVIDIA converge in 2022. The biggest datacenter hardware refresh in a decade.

Only 1% of Software Benefits from SIMD Instructions

Analysis of 2,000 Linux and macOS binaries reveals less than 1% of instructions use SIMD, despite CPUs dedicating massive die area to vector processing.

Artsakh Must Be Independent 🗺️

A data-driven argument for why Artsakh must be independent — examining historical claims, demographic evidence, and the indisputable pattern of ethnic violence.

The 7 Sins of Turkish Autocracy 🇹🇷

Seven stories tracing Turkey's path from secular reforms to neo-Ottoman ambitions — examining crimes unpunished, allies betrayed, and the dangerous pattern repeating today.

Armenia, Azerbaijan, Turkey. Who's the Aggressor? ⚔️

Three countries, vastly different resources and freedoms. The numbers reveal who's really threatening peace in the region.

Come to Armenia 🇦🇲

Why Armenia is an emerging tech hub worth your attention — from startup tax incentives to ancient history, discover what makes this small nation punch above its weight.

Positive Outlook on the COVID-19 Crisis 😷

The pandemic revealed our weaknesses, but also unprecedented opportunities — from personal growth to financial markets, here's why optimism isn't naive.

Building AI Safely

A 2018 conversation about AI safety, infrastructure, and the real dangers of weak AI in the hands of organizations optimizing for power over people.

What's Wrong with WWDC 2016 Keynote?

An iOS developer's critique of Apple's 2016 keynote — when emojis and animations overshadowed the innovations developers actually needed.

HackerNews Favorites

I often post on Reddit and HackerNews, with the latter being a surprisingly well-balanced platform for technical discussions with less personal bias. Here are some of my favorite publications that made headlines on HackerNews: Up to 100x Faster FastAPI with simdjson and io_uring on Linux 5.19 . Beating OpenAI CLIP with 100x less data and compute . Less Slow C++ . Full Unicode Search at 50× ICU…

My Open Software

All of my software is hosted on GitHub, mostly under the Apache-2.0 permissive license. Free for commercial and non-commercial use, modification, and distribution. Major Projects USearch - a universal search engine powering many databases, AI labs, and experiments in Natural Sciences. Compact C++ core with 10+ language bindings — 10–100× faster than Meta FAISS for vector search and far beyond…

Recordings & Talks

Most materials are in English unless literally flagged otherwise. The absolute majority is on the subjects of Systems Design, Computer Science, and Artificial Intelligence. The 🗣️ talking head links aren’t technical, and in the ones with a 👯‍♂️ - I am just a wingman supporting another speaker. 2025 PyTorch & Lightning AI Meetup: Matrix Multiplication Assembly Instructions. London, UK.…