Log inSign up
Log inSign up
Harshay Shah
24 posts
@harshays_

Harshay Shah

@harshays_
research @googledeepmind
New York
harshay.me
Joined May 2012
446
Following
646
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @harshays_
    Harshay Shah
    @harshays_
    Jan 29, 2025
    MoEs provide two knobs for scaling: model size (total params) + FLOPs-per-token (via active params). What’s the right scaling strategy? And how does it depend on the pretraining budget? Our work introduces sparsity-aware scaling laws for MoE LMs to tackle these questions! 🧵👇
    @samira_abnar
    Samira Abnar
    @samira_abnar
    Jan 28, 2025
    🚨 One question that has always intrigued me is the role of different ways to increase a model's capacity: parameters, parallelizable compute, or sequential compute? We explored this through the lens of MoEs:
    1
  • @harshays_
    Harshay Shah
    @harshays_
    Apr 19, 2024
    New work with @andrew_ilyas and @aleks_madry on tracing predictions back to individual components (conv filters, attn heads) in the model! Paper: arxiv.org/abs/2404.11534 Thread: 👇
    @aleks_madry
    Aleksander Madry
    @aleks_madry
    Apr 18, 2024
    How do model components (conv filters, attn heads) collectively transform examples into predictions? Is it possible to somehow dissect how *every* model component contributes to a prediction? w/ @harshays_ @andrewilyas, we introduce a framework for tackling this question!
    1
  • @harshays_
    Harshay Shah
    @harshays_
    Jul 26, 2023
    If you are at #ICML2023 today, check out our work on ModelDiff, a model-agnostic framework for pinpointing differences between any two (supervised) learning algorithms! Poster: #407 at 2pm (Wednesday) Paper: icml.cc/virtual/2023/p… w/ @smsampark @andrew_ilyas @aleks_madry
  • @harshays_
    Harshay Shah
    @harshays_
    Dec 7, 2021
    Do input gradients highlight discriminative and task-relevant features? Our #NeurIPS2021 paper takes a three-pronged approach to evaluate the fidelity of input gradient attributions. Poster: session 3, spot C0 Paper: bit.ly/3EzdvyH with @jainprateek_ and @PNetrapalli
  • @harshays_
    Harshay Shah
    @harshays_
    Dec 9, 2020
    Neural nets can generalize well on test data, but often lack robustness to distributional shifts & adversarial attacks. Our #NeurIPS2020 paper on simplicity bias sheds light on this phenomenon. Poster: session #4, town A2, spot C0, 12pm ET today! Paper: bit.ly/39RXDel
    Prateek Jain and Praneeth Netrapalli