Join freshers getting daily off-campus drives, direct apply links & remote internship updates.
100% Free · No spam · Instant direct-apply links only
Experience / Eligibility
B.E / B.Tech / M.Tech
Salary
Not Disclosed / As per Industry Standards
Location
Bangalore, Karnataka, India
Suitable For
College graduates, entry-level candidates, and students matching: B.E / B.Tech / M.Tech.
Key Skills to Prepare
Focus on Python, Machine Learning, Artificial Intelligence (AI), PyTorch.
Pre-train STT models from scratch (Conformer/Zipformer encoders, CTC/RNN-T/TDT decoders, self-supervised pre-training like wav2vec 2.0/HuBERT/BEST-RQ), then fine-tune them for Indian languages, Hinglish and BFSI vocabulary
Train TTS models from scratch (flow-matching, diffusion, neural codec LMs, vocoders) and extend them to new voices, languages and speaking styles
Build speech-to-speech components from scratch: audio tokenizers and codecs, speech encoders connected to LLMs, streaming decoders
Turn papers and new ideas into working PyTorch code: write the model, tokenizer and training loop yourself, and train from random initialisation
Build data pipelines: cleaning, segmentation, forced alignment, pseudo-labelling, augmentation (noise, codecs, 8 kHz telephony simulation)
Maintain evaluation suites: WER/CER, entity error rate, code-switch accuracy, MOS and speaker similarity, latency (time to first token / byte)
Optimise models for serving: quantisation, streaming chunking, ONNX/TensorRT export, batching on GPUs
Analyse production errors, turn them into test sets, and close the loop with the platform team
Write clear experiment reports and share findings in weekly research reviews
What you bring
1-3 years in ML, with at least one real project in speech or audio (industry, research lab, or a strong thesis)
B.Tech/M.Tech/MS in CS, EE or a related field; a PhD is not required
Strong Python and PyTorch; able to implement a model architecture and training loop from scratch, not just call APIs or fine-tune checkpoints
Working knowledge of speech basics: spectrograms, MFCC/mel features, CTC, attention-based seq2seq, vocoders
Hands-on with at least one of: NeMo, ESPnet, Hugging Face Transformers/Audio, fairseq, Coqui, k2/icefall
Has trained at least one model from random initialisation (speech, audio or language), including multi-GPU jobs, mixed precision and experiment tracking (W&B, MLflow)
Careful with evaluation: you know why a WER number can mislead, and you check your data
Speak or understand Hindi or another Indian language
Publications, Kaggle/benchmark results, or open-source contributions in speech
Exposure to real-time audio (WebRTC, SIP, streaming inference) or telephony audio
Familiarity with LLM fine-tuning (LoRA, SFT) or audio-language models
What success looks like in 6 months
Contributed to a pre-trained-from-scratch model that shipped to production with a measured gain on a customer-relevant test set
Owns a piece of the evaluation or data pipeline that others rely on
Runs experiments independently and reports results the team can trust
Essential Python interview questions asked by product startups and service-based IT companies during fresher campus and off-campus placements.
Technical PrepPrepare for your entry-level Python developer interview with 50 high-frequency questions covering data types, memory management, list comprehensions, decorators, and DSA.