Hey, I'm Aditya - a Computer Science (Data Science) student and AI-focused builder with hands-on experience in machine learning, software engineering, and data workflows. I've taken technical projects from data collection and preprocessing through automation, validation, and working implementations, including building the core training dataset for a production text-to-speech model. I work primarily in Python and Rust, with a strong interest in AI products, rapid experimentation, automation, and practical problem solving. Currently open to internships, research opportunities, and collaborations.
Built and curated the core training dataset for EvoTalk, the company's text-to-speech (TTS) model, later
released on Hugging Face and GitHub. Sourced raw audio-text pairs and defined data quality criteria to
filter noisy, misaligned, or clipping samples before ingestion. Designed and ran automated preprocessing
pipelines for text normalization, forced alignment, phoneme mapping, and deduplication, and performed
quality assurance audits on transcription accuracy and audio integrity. Collaborated with ML researchers
on end-to-end dataset preparation, validation, and release workflows.
A Python and Rust tokenizer that generates sparse token representations, prioritizing infrequent but
significant tokens for efficient handling of large vocabularies. Designed a configurable sparsity
threshold to control the trade-off between memory footprint and vocabulary coverage, and benchmarked the
Rust implementation against Python on large-scale text corpora, evaluating runtime performance against
development speed.
An ML pipeline to analyze scanned historical documents, estimate authenticity, and detect potential forgery
anomalies. Extracted handwriting style features and linguistic patterns, comparing them against verified
authentic documents to flag stylistic and linguistic anomalies, and built an interpretability dashboard
with authenticity confidence scores and regional heatmaps for historians and archivists.