Sameer Banchhor

Sameer Banchhor

AI/ML Researcher • Data Scientist • NLP & Speech Systems for Low-Resource Languages

I design and build end-to-end machine learning systems focused on educational data mining, predictive modeling, and speech/conversational AI for under-represented languages. From processing 248,000+ academic records to deploying fine-tuned speech models in rural settings, my focus is bridging the gap between raw data and real-world impact.

"Building AI that speaks every language — including the ones the world forgot."
Exam Records Processed 248,539 DURG-EduAI Academic Dataset
Classification F1 Score > 0.99 R² ~ 0.996 SGPA Regression
TTS Speaker Language Impact 18M+ Chhattisgarhi VITS Model
Multi-Task Model Outputs 5 Targets Academic Risk & Early Warnings

Research & Featured Projects

Machine Learning Systems, Low-Resource NLP, and Field Deployments

DURG-EduAI — Multi-Task Student Performance Prediction

Hemchand Yadav University, Durg • Educational Data Mining
Aug 2025 – Feb 2026
Large-scale machine learning framework designed for academic risk analysis, trained on 248,539 student examination records spanning 2016 to 2025.
Achieved classification F1 > 0.99 and regression R² ~ 0.996 on structured academic records.
Proposed a novel 6-signal proxy-label methodology for dropout risk prediction without longitudinal tracking.
Delivers 5 simultaneous output predictions per student (SGPA, ATKT status, dropout risk, benchmarks, alerts).
XGBoost LightGBM Scikit-learn Pandas Python
DURG-EduAI is a large-scale machine learning framework designed for academic risk analysis, trained on 248,539 student examination records spanning 2016 to 2025 at Hemchand Yadav University, Durg. The system delivers five simultaneous predictions per student: 1. SGPA Regression 2. Pass/Fail/ATKT Classification 3. Dropout Risk Detection 4. Subject-Level Benchmarking 5. Early Warning Alerts Findings were presented to university leadership; models and datasets were prepared for open platform publication.

Chhattisgarhi Text-to-Speech (VITS Deep Neural Network)

Independent Research • Speech AI
2025 – 2026
One of the pioneer Text-to-Speech (TTS) architectures developed for Chhattisgarhi, a low-resource language spoken by over 18 million people in Central India.
Trained a VITS (Variational Inference with adversarial learning) deep neural network.
Constructed an end-to-end pipeline from Devanagari text normalization to raw audio waveform synthesis.
Overcame dataset scarcity and handled regional phonetic/tonal variations.
VITS PyTorch HuggingFace Phoneme Processing
One of the first TTS systems built for Chhattisgarhi, a low-resource language spoken by over 18 million people in Central India. Key Highlights: - Trained a VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech) deep neural network. - Addressed core low-resource challenges: limited training data, dialect variation, and tonal pronunciation. - Designed a complete pipeline from Devanagari text normalization to audio waveform generation.

Chhattisgarhi Conversational AI (Field Deployment)

Kanker, Chhattisgarh • Field AI & Dialect Modeling
Mar 2025 – Jul 2025
A fine-tuned conversational agent developed to support local agricultural communities with queries on farming tools, seeds, and seasonal practices.
Fine-tuned LLM architectures (Gemini API) on regional dialects and colloquial speech patterns.
Deployed in real-world rural environments with direct iterative feedback loops from farmer groups.
Recognized and compensated by local government for community technology impact.
Gemini API Fine-tuning Transformers NLP
A chatbot fine-tuned on a custom Chhattisgarhi corpus to support local farming communities with queries on seeds, tools, and agricultural practices. Key Highlights: - Fine-tuned Google Gemini API on regional dialects and colloquial speech patterns. - Deployed in a real-world rural setting in Kanker, Chhattisgarh with direct feedback loops from farmer groups. - Recognized and compensated by local government for community technological impact.

DU-Analysis — Academic Data Mining Archive

Public Codebase • Data Engineering
2025
Exploratory data analysis workspace and data preprocessing pipelines built to clean, ingest, and structure raw unstructured academic records into clean tabular formats.
Python Pandas ETL Pipelines

Technical Expertise

Tools, Frameworks & Languages
Languages
Python SQL JavaScript HTML/CSS
Machine Learning & AI
Scikit-learn XGBoost LightGBM PyTorch HuggingFace Gemini API
Speech & NLP
VITS Architecture Transformers Text Normalization Phoneme Tokenization
Data Science & Engineering
Pandas NumPy Matplotlib Seaborn Git ETL Pipelines

Education & Background

Academic Degrees and Professional Certifications

M.Sc. in Computer Science

Hemchand Yadav University, Durg
8.21 SGPA (Sem III)
2024 – 2026

B.Sc. in Computer Science

Kalyan PG College, Bhilai
79%
2021 – 2023

Google Data Analytics Professional Certificate

Coursera
Completed
2026

IBM Data Science Professional Certificate

Coursera
Completed
2026