About

👋 Hi, I’m Akash, an applied researcher/engineer with experience in speech, audio (at Microsoft), and most recently multi-modal document understanding and retrieval (at Contextual AI). This incidentally completes the trio of audio, vision & text AI multimodality. :)

I’m currently on a sabbatical. After moving to the USA for grad school ~10 years ago, I decided to take a break to reflect, recharge and tinker before setting sail again. More on this here shortly!

Work

Contextual AI

[2024-25]

Wrangled millions of pages to land the first $ millions in enterprise contracts :)

  • System development (0→1)
    • Designed core multimodal document understanding (parsing + ETL index) used by retrieval agents for all customers in production
    • Ingesting O(billion) tokens of multimodal documents with complex layouts, tables, graphics, metadata
    • Shipped as a distributed async service with multiple stages (CPU, GPU and VLM API). I owned system quality and core modules, and co-led architecture with a platform lead
  • Applied research: harness and eval design, token-efficient retrieval
  • Tech Lead Manager
    • PoC with Product/GTM/Marketing across company; Interviewed candidates, led team of 3
    • Critical in landing company’s first multi-million $ enterprise contract with Qualcomm

Microsoft

[2018-23]

Fun fact: ~6M hours of monthly traffic equals 1 *year* of conversations transcribed per hour!

  • Model development: state-of-art transcription designed for scale [O(10 million) hrs/mo]
  • Applied research: diarized multi-speaker multi-mic transcription
    • Shipped diarized in-conference room transcription device covered by The Verge
    • Lead contributor: ASR training recipes, evaluation metrics, cross-system error analysis
  • Research engineering: data pipelines, distributed training, inference
    • Sped up O(1e20) FLOP training on low-cost V100 GPUs to run in <1 week
    • Fixed inference bottlenecks leveraging NVIDIA/ONNX profiling tools, saving $ millions
  • Other Links:

‘Graduated’ as one of the few non-speech-PhD senior members on the team :)

For more details, see my resume.

Misc

Open source

Other

  • [2025/26] Peer reviewer for ICASSP conference, TMLR journal, and NeurIPS, CVPR workshops.
  • [2016/17] Wrote case studies on the music streaming industry while studying business/tech strategy at Stanford MS&E.
  • [2014] Organized (at the time) Chennai’s largest EDM gig - with 5k+ attendees, during my undergrad at IIT Madras/Chennai.