Arnab Sen Sharma · X (formerly Twitter)

user avatar

Ph.D. student

@KhouryCollege

, working to make LLMs interpretable

Boston, MA

Joined September 2022

  • Pinned

    user avatar

    How can a language model find the veggies in a menu? New pre-print where we investigate the internal mechanisms of LLMs when filtering on a list of options. Spoiler: turns out LLMs use strategies surprisingly similar to functional programming (think "filter" from python)! 🧵

  • user avatar

    Super excited to be attending

    @iclr_conf

    in Rio. Stop by our poster tomorrow morning (10:30am - 1:00pm) in Pavilion 4 (P4-#4001) to know about list-processing mechanisms in LMs. DMs are open. Please reach out if you want to meet up!

  • user avatar

    Attending

    @COLM_conf

    in Philadelphia! Drop by our poster on Wednesday morning (Session 5, 11 am - 1 pm, #8) Would love to catch up and chat about interpretability. Give me a DM!

  • user avatar

    Just reached Vienna to attend ICLR! Stop by our poster session tomorrow (10:45 am, Hall B, #131). Would love to chat with people about interpretability and AI alignment. DMs are open!

  • user avatar

    How does Mamba store knowledge? Is it very different from transformers? New pre-print with

    @diatkinson

    and

    @davidbau

    , where we investigate the mechanisms of factual recall within Mamba.

Read the original on x.com ↗