RSS Amplifier

Blog

Low Latency Trading Insights

Fresh insights on low latency and high frequency trading. Mostly new relevant technology, C++, FPGA, Verilog, tutorials, short notes and of course, rants. Problem subscribing? Try https://lucisqr.lemonsqueezy.com/buy/c78e2dd6-81c0-44cd-bfa3-13db7158a050

lucisqr.substack.comSource feed ↗15 posts

Live Last read · last published · next check

Latest posts

Saves to your Listen queue, to pick up on another day or another device.

Your AVX-512 GCD Isn't Slow. Your Benchmark Is.

A reader sent me a small experiment.

The Most Useful Pointer the STL Never Shipped

Lay the three side by side: unique_ptr is eight bytes and cannot share.

Armed and dangerous: std::shared_ptr<T> considered harmful

A few years ago I was looking at some performance graphs and I got surprised to see that I got a huge performance hit when replaced my homegrown smart pointer with the STL std::shared_ptr and kept wondering why since these should be mostly equivalent.

Better Descriptive Statistics for Low-Latency Profiling

Many C++ developers have a pretty good idea about measuring micro-events - but if they don’t, we wrote another article just about that a few months ago: Microbenchmarking is tricky!

BOLT After PGO: A 4% Win That Used to Be a 30% Loss!

If you liked it, here is another one!

Your distro already PGO'd your compiler. So what does super-gcc actually give you?

Compile time is one of the bottlenecks HFT engineers notice every day.

Programmable Last Level Cache: From Theory to Production (Part 2)

In Part 1 we covered why shared L3 cache is the silent performance killer in multi-workload systems, and how Intel's Cache Allocation Technology (CAT) and AMD's Platform QoS let you carve up the LLC into isolated partitions using Classes of Service (CLOS) and Capacity Bitmasks (CBM).

The One Honest Use of C++20 Concepts

Someone on the desk asks whether we should be putting concepts on the hot path.

LLVM Benchmarked Swiss Tables and Kept Its Own Hash Map

LLVM is one of the most performance-obsessed C++ codebases on the planet.

The 16-Byte Rule, the 64-Byte Rule, and the 128-Byte Rule Are Three Different Compilers' Opinions

I put every struct-size rule C++ engineers memorize on five machines and eleven toolchains.

A Hash Table Smaller Than Its Data Is Real.

The Version You Can Copy Is Not.

Play

std::hive Erases 36x Faster Than std::list. Your Order Book Will See 1.8x. Switch Anyway.

Internally a hive is a linked list of blocks with geometrically growing capacities.

The Magic Ring Buffer: A Beautiful Trick That's Mostly Not Worth It

Map the same block of memory twice, right next to itself, and a ring buffer’s wrap-around just… vanishes.

Your Hash Table Doesn't Need Trees. It Needs a Resize.

A paper that went up on arXiv two days ago reports that inserting 500,000 keys into a deliberately overloaded C hash table takes 271.5 seconds with linked-list buckets and 121.2 seconds with “hybrid-batch” buckets.

Zen 6, Diamond Rapids, Vera: What 2026 Silicon Buys the Hot Path

AMD formally launches Zen 6 on Tuesday, July 22, leading with the server part, and by Wednesday the coverage will have mixed the keynote numbers, the leaks, and two-year-old news into one undifferentiated pile.