Achyuthan-S
Achyuthan Sivasankar

Now

MS CS at NYU, graduating Fall 2027. Research assistant in Prof. Anna Choromanska's lab. Targeting PhD programs in core ML.

Adaptive computation · MoE systems · Efficient architectures

I build and study adaptive routing in sparse neural architectures, from function-basis routing in KAN layers to expert-collapse dynamics in large MoE language models.

+1,722 FSD lead steps before grokking
+6.8% CIFAR-100 over MLP (KAN-Multi)
99.67% SWELL-KW stress detection
2× AutoMoE parameter efficiency

Research interests

Mechanistic interpretability Grokking MoE architectures KAN layers Adaptive computation Self-supervised learning

Education & affiliations

  1. MS Computer Science

    New York University

    Graduating Fall 2027 · NYC

  2. Research Assistant, Prof. Anna Choromanska's Lab

    NYU · AD-LiST-JEPA · Waymo Open Dataset

    May 2026 - Present

  3. Research Intern, Prof. Sunil Chandran

    IISc Bangalore · GNN routing · EEG / BCI

    Jun - Aug 2024

  4. AI Research Intern

    National University of Singapore

    Dec 2023 - Feb 2024

Contribution activity and merged upstream work in large training and inference frameworks.

@Achyuthan-S

ML research · inference runtimes · MoE systems

View profile
… Public repos
… Followers
Achievements

Contribution activity

Notable upstream contributions

DeepSpeed · vLLM · NVIDIA NeMo · Megatron-LM · Megatron-Bridge

Featured

Subsystems  ·  where the work lands

The stack, top to bottom 19 merged
Inference serving & KV cache Requests in, tokens out 3
Weight tying & model loading Checkpoint config to live model 6
Checkpoint shard geometry Where the bytes live across ranks 5
Distributed training Gradients across ranks 2
Data loading & resume Replaying the right samples 1
Build & runtime Import, compile, ship 2

Every merge  ·  19 total

Pull requestFilesReviewMerged
Restore a shard from the map the layer published deepspeedai/DeepSpeed #8622 4 5 Sep 29 Describe the GPTBigCode and Yuan shared-QK layouts deepspeedai/DeepSpeed #8575 4 8 Sep 20 Emit affine maps from AutoTP layers deepspeedai/DeepSpeed #8519 4 8 Sep 16 Carry the affine scale on the replicated map, not the split deepspeedai/DeepSpeed #8477 2 3 Sep 11 Describe universal checkpoint shards as affine maps deepspeedai/DeepSpeed #8385 8 26 Sep 10 Support HFSDP deferred DP-outer gradient reduction NVIDIA/Megatron-LM #5972 2 52 Aug 15 Disable cross-layer KV blocks for per-token-head quantized KV cache vllm-project/vllm #49226 1 12 Jul 26 Support HSDP deferred DP-outer gradient reduction NVIDIA/Megatron-LM #5743 4 94 Jul 20 Single tie_word_embeddings guard via TieSupport NVIDIA-NeMo/Automodel #2998 63 23 Jul 16 Complete tie_word_embeddings guards for remaining model families NVIDIA-NeMo/Automodel #2896 8 6 Jul 8 Make finetuning batch sampler epoch-aware on checkpoint resume NVIDIA-NeMo/Megatron-Bridge #4601 2 4 Jul 6 GPT-OSS: recover raw tail when Harmony parser ends non-terminal vllm-project/vllm #47379 4 3 Jul 4 GPT-OSS: return raw output when Harmony parser ends non-terminal vllm-project/vllm #47062 2 15 Jul 1 Reject tie_word_embeddings=True on separate-head model families NVIDIA-NeMo/Automodel #2805 28 17 Jul 1 Avoid CUDA context initialization during import-time op compatibility checks deepspeedai/DeepSpeed #8078 14 10 Jun 29 Resolve tie_word_embeddings top-level-first to match Hugging Face tying NVIDIA-NeMo/Automodel #2732 3 9 Jun 25 Cherry-pick #2601 into r0.5.0 release branch NVIDIA-NeMo/Automodel #2709 2 3 Jun 22 Re-tie lm_head to active embed_tokens on Gemma4 MoE path NVIDIA-NeMo/Automodel #2601 2 20 Jun 22 Fix nightly Docker ImportError: AnthropicOutputConfig vllm-project/vllm #44795 2 12 Jun 14

Opens your email client with your message pre-filled.