Saoud Rizwan · X (formerly Twitter)

Saoud Rizwan

Cline

1,175

posts

Saoud Rizwan profile banner

user avatar

Cline

@sdrzn

San Francisco

Joined March 2011

  • Pinned

    user avatar

    kimi k3's press release says "built on Kimi Delta Attention" a ton so i read their paper on it, incredible stuff. "normal attention" is like reading a whole book before answering each question on a test. each new token looks back at ALL previous tokens, so compute scales with

  • user avatar

    databricks switched from opus to glm, which their evals found had 92% of the pass rate at 1/4th the cost. ...their evals on their own proprietary codebases, not on public benchmarks that models could have hillclimbed. the chinese benchmaxxing accusations are officially cope.

    user avatar

    Today

    @databricks

    we're publishing a detailed analysis of techniques we used to drastically reduce our internal AI spend while aggressively growing adoption. Savings come from layering in several techniques, which combine to drive unit costs down as much as 90% in some scenarios.

  • user avatar

    Incredibly grateful to stand with all the leaders in this space rallying around open weights. Thank you to everyone that has contributed to open source, whether that's a PR, a research paper, or a signature.

    user avatar

    Cline has signed the Open Weights letter, and we're proud to be a part of this movement. Millions of developers use open weights in Cline because they're cheaper, address data privacy concerns, or meet regulatory restrictions. To celebrate, we're making GLM 5.2 free in Cline!

  • user avatar

    > huggingface gets breached and forced to use glm bc of fable refusals > kimi k3 discovers and reports countless 0days in redis a week later glad open weights can't be gatekept behind refusals, current sota half baked nerfing is only hurting defenders.

  • user avatar

    cloud was effectively commoditized in march 2014 when google cut compute prices 32%, and aws was forced to match within days with its 42nd price cut at the time. inference is speedrunning the same story, when dollars are involved markets are brutally efficient.

    user avatar

    Kimi costs ~3-12x cheaper than Fable, but how much more could you save hosting it yourself? We ran the numbers on Cline’s production traffic, and the results: ~10% savings, 25%+ with time-of-day autoscaling (but this only works if over $500K/year of spend) We predict

Read the original on x.com ↗