Skip to content
View yellatp's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report yellatp

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
yellatp/README.md
Pavan Yellathakota

Hey DataGeeks, I'm Pavan 👋

I'm an AI/ML Engineer who understands the data world from the ground up. I started in customer and consumer analytics as a Data Analyst, moved into supply chain analytics, inventory management, and seasonal and event-based demand forecasting as a Data Scientist, and now build and deploy production ML models and Generative AI systems. Along the way I consulted for 10+ local businesses through Clarkson University, including HAVK Mladost, co-founded my own venture, Alphonso AI (alphonso.app), and launched onlynerds.win, a resources and job search platform for job seekers. I'm currently an AI/ML Engineer at USK Systems, consulting and building products for Fortune 100 and Fortune 500 clients as well as promising AI startups.

AI/ML Engineer | Data Scientist | Generative AI

📍 Location : Austin, TX, USA
📞 Mobile : +1 (929) 278-4589
✉️ Email : pavan.yellathakota.ds@gmail.com
Linkedin : https://linkedin.com/in/yellatp
GitHub :   https://github.com/yellatp
Website :   onlynerds.win



👨‍💻 Professional Summary

AI/ML Engineer with 4+ years of experience across the full arc of the data career path. I began in customer and consumer analytics as a Data Analyst, grew into a Data Scientist role building supply chain analytics, inventory management systems, and seasonal and event-based forecasting models, and now work as an AI/ML Engineer developing and deploying production ML and Generative AI systems. I have consulted for 10+ local businesses through Clarkson University, co-founded my own venture, Alphonso AI, and launched onlynerds.win, a resources and job search platform for job seekers. I currently work as an AI/ML Engineer at USK Systems, consulting and building products for Fortune 100 and Fortune 500 clients as well as promising AI startups.


🛠️ Technical Skills

Domain Stack
Languages & Databases Python Pandas NumPy Scikit-learn SQL PySpark PostgreSQL R TypeScript
AWS Cloud Data AWS S3 Athena Glue SageMaker Lambda Redshift
ML Frameworks XGBoost Hugging Face OpenAI NLTK Spacy
Tools & Visualization Tableau QuickSight Power BI Excel Git

💼 Professional Experience

USK Systems | AI/ML Engineer, Consultant

Remote | 2026 – Present

  • Work as a consulting AI/ML Engineer, designing and building AI products for Fortune 100 and Fortune 500 clients as well as early-stage AI startups.
  • Partner directly with client teams to scope ML and Generative AI solutions, moving from problem definition to a deployed, production-ready system.
  • Apply the same end-to-end approach used across my own ventures, including data pipeline design, model development, and deployment, to client engagements spanning multiple industries and company sizes.

Alphonso AI, backed by Shipley Center for Innovation | Co-Founder & Founding ML Engineer

Potsdam, NY | Jul 2025 – Present

  • Joined as a founding engineer with no existing bridge between the company's Java-based core services and its Python-native ML workloads. Designed and built a 0→1 backend ecosystem using FastAPI and PostgreSQL that now runs in production as the scalable microservices layer connecting both sides of the stack.
  • Established and manage production services on DigitalOcean VPS to keep infrastructure spend lean for an early-stage startup, and introduced Docker-based containerization so R&D and production environments stay in full parity.
  • Engineered a multi-model Text-to-Query (TTQ) system that integrates Gemini (Vertex AI) and DeepSeek APIs, giving users dynamic, prompt-driven semantic search across a high-dimensional talent database.
  • Implemented a multi-stage retrieval pipeline using pgvector for Approximate Nearest Neighbor (ANN) search combined with CUDA-accelerated cross-encoders for re-ranking, targeting a 38% improvement in Precision@N over the prior search approach.
  • Redesigned the recommendation engine, moving matching logic from generic role-based matching to a sector-specific ranking system powered by vectorized embeddings, improving candidate-to-company fit.
  • Delivered a generative team-composition module that converts plain-language product descriptions into detailed technical requirements and concrete candidate matches, closing a real usability gap for non-technical founder users.
  • Led relational schema normalization and API contract design, and directed early research into Model Context Protocol (MCP) for agentic, self-correcting database interactions.

Key Technologies Used
Python FastAPI PostgreSQL Docker pgvector HuggingFace Gemini Vertex AI

Student Managed Investment Fund, Clarkson University | Graduate Quantitative Researcher

Potsdam, NY | Sep 2024 – Apr 2025

  • Contributed core research and risk analysis behind a $650K real-capital portfolio, helping the fund deliver a 51% total return and outperform the S&P 500 benchmark by 26% (2,600 bps).
  • Developed a sentiment analysis engine that scrapes Reddit and YouTube and scores sentiment using a BERT-based model, giving the team a live sentiment overlay to validate fundamental buy signals before committing capital.
  • Automated the extraction of financial statements from SEC EDGAR using Python and Vertex AI, cutting data collection time by 80% for the analyst team.
  • Built Monte Carlo simulations and risk-parity models to stress-test overweight positions and quantify drawdown risk on high-conviction trades before they were placed.

Key Technologies Used
Python BERT HuggingFace Vertex AI Pandas

HAVK Mladost (Elite Athletics Club) | Graduate Data Science Consultant

Potsdam, NY | Oct 2023 – May 2025

  • Established a centralized data lake on AWS S3, migrating legacy records into a queryable cloud environment and cutting data retrieval latency by 40%.
  • Developed PySpark ETL jobs on AWS Glue to process over 1M cross-channel events, applying partition pruning to reduce query costs and improve speed.
  • Applied uplift modeling and behavioral clustering to identify high-value fan segments, directly informing marketing spend and merchandise strategy.
  • Delivered backend services with FastAPI and the interactive dashboards on top of them, providing real-time performance insights to World Championship coaches.

Key Technologies Used
AWS S3 Glue PySpark FastAPI Python

eAppSys Limited | Business Data Analyst

Hyderabad, India | Jul 2022 – Dec 2022

  • Built demand forecasting models (Prophet/SARIMAX) for over 1,500 SKUs, incorporating exogenous variables such as holidays and promotions to improve forecast accuracy (MAPE) by 15%.
  • Designed and deployed automated KPI dashboards in Oracle Analytics Cloud (OAC), saving the procurement team 12+ hours per week of manual reporting.
  • Implemented GxP-compliant ML workflows on Oracle Cloud Infrastructure (OCI) with real-time alerting, achieving 99.9% uptime for critical inventory monitoring.

Key Technologies Used
Python Prophet SARIMAX Oracle OCI

Kantar GDC India | Data Analyst

Pune, India | Sep 2021 – May 2022

  • Built automated data pipelines for Tracker and Syndicated Research projects using Python and PySpark, integrating over 10M survey records from 30+ sources and cutting processing latency by 30%.
  • Developed sampling approaches and statistical significance testing to ensure data representativeness across Middle East and Central Africa markets.
  • Built regression models supporting recurring monthly and quarterly client tracking, delivering actionable insights for 10+ FMCG and Telecom clients.

Key Technologies Used
Python PySpark Pandas SQL


🚀 Independent Ventures

Alphonso AI — Co-Founder and Founding ML Engineer. I helped start Alphonso AI, backed by the Shipley Center for Innovation, and lead the ML and backend engineering side of the company, covering everything from the 0→1 backend architecture to the search and recommendation systems described in the experience section above. Full details on this work are in Professional Experience.

onlynerds.win — Only Nerds, an ATS Intelligence Platform I conceived, designed, and built on my own, outside of any employer or client engagement. It maps 115,008+ unique company records to the applicant tracking system (ATS) each one runs on, including Greenhouse, Ashby, Lever, BambooHR, iCIMS, Workable, join.com, Personio, and 20+ more, so job seekers can identify a company's ATS and reach its job board in one click. The site runs on Astro and TypeScript, and I handle everything end to end: product direction, data pipeline, frontend build, and deployment.

  • ATS Lookup — Search 115K+ companies to see which ATS platform they use, with a direct link to their job board.
  • Company Directory — Browse H1B sponsors (2,200+), private market firms (4,200+), and VC portfolios (88 funds, 25 accelerators).
  • Google Dorking Engine — An interactive query builder that targets specific roles, ATS platforms, experience levels, and locations across Google, Bing, and DuckDuckGo.
  • Region Intelligence — Filter companies by region (US-CA, EU, ASIA, LATAM, AU-NZ, AFRICA) based on ATS headquarters and location data.
  • Parquet-powered — All company data, covering 4.2M+ jobs and 115K companies, is stored in Apache Parquet and loaded client-side via WASM, keeping the platform fast without a heavy backend.

🏗️ Notable Projects

Each project below reflects the actual scope of the underlying repository, including the problem it addresses, the sector it applies to, and the technical approach behind it.

Flagship Projects

Detoxify Telugu
A fine-tuned BERT-based language model for hate speech detection in Telugu and Tenglish (Telugu written in Latin script). Large general-purpose language models are trained overwhelmingly on high-resource languages and routinely miss dialect-specific slang and mixed-script text, which lets harmful content slip past moderation in regional markets. This project fine-tunes a transformer model specifically on Telugu and Tenglish text to close that gap.

  • Relevant sectors: Social platforms, regional content moderation, LLM safety and alignment
  • Stack: Python, PyTorch, Hugging Face Transformers, BERT

BingeMax Recommendation Engine
An AI-powered movie recommendation system that combines content-based filtering, collaborative filtering, and cosine similarity scoring to generate personalized suggestions, served through a Streamlit interface backed by a FastAPI service layer.

  • Relevant sectors: Streaming platforms, e-commerce, adtech
  • Stack: Python, Streamlit, FastAPI, Scikit-learn

Fintech Sales GAP Analysis
A Python-based analytics project built for fintech sales teams to quantify performance gaps across sales representatives, regions, and product lines, translating the findings into concrete, data-driven training recommendations for underperforming segments.

  • Relevant sectors: Fintech, banking, enterprise revenue operations
  • Stack: Python, Pandas, Scikit-learn, PostgreSQL, Plotly

KonnectR
A full-stack web application that connects academia and industry, letting students, professors, and professionals collaborate in one place. It includes an asynchronous chat system, listings for jobs and research opportunities, and role-based user management, and was later extended into a second build, KonnectR_flask_fullstack_app, adding analytics and posting modules on top of the same concept.

  • Relevant sectors: EdTech, academic research networks, professional knowledge sharing
  • Stack: Python, Flask, PostgreSQL, async messaging, REST APIs

CUDA vs CPU Showdown
A hardware-aware performance benchmarking study comparing GPU-accelerated data processing against traditional CPU workflows, using RAPIDS (cuDF) and DuckDB across datasets of over 1M rows on an NVIDIA GTX 1650. The results document a 3x to 10x throughput improvement from GPU acceleration, making the case for hardware-aware pipeline design in data engineering work.

  • Relevant sectors: Data engineering, ML infrastructure, high-throughput analytics
  • Stack: Python, RAPIDS (cuDF), DuckDB, CUDA

XIFTY
An edge-native, desktop-first intelligence suite built to scout, analyze, and manage content creators at scale, designed for fast local performance rather than a purely cloud-dependent architecture.

  • Relevant sectors: Creator economy, influencer analytics, martech
  • Stack: TypeScript

Additional Projects

AtmosDB — A unified SDK that brings Cloudflare D1, Vectorize, and R2 together behind one interface, aimed at building global, AI-native backends at the edge without the usual dashboard overhead. (TypeScript)

Customer Acquisition Cost Analysis — An in-depth analysis of customer acquisition cost across marketing channels for 2023, evaluating CAC, conversion efficiency, and spend effectiveness to guide marketing budget decisions. (Python, Pandas)

Text Analysis using NLP and LDA — An NLP project covering sentiment analysis, named entity recognition, and topic modeling with LDA, applied to unstructured text corpora. (Python, NLTK, Gensim)

Fake News Classifier — A machine learning system that classifies news articles as fake or real using Naive Bayes and SVM models, with a full text preprocessing and evaluation pipeline. (Python, Scikit-learn)

GenZ Career Preferences Report — A data-driven study of career preferences among Gen Z professionals, based on survey responses from 230+ participants, presented through an interactive dashboard. (Python, Pandas, Plotly)

Content Strategy Analysis: Netflix — An analysis of Netflix's 2023 viewership trends, examining content performance, audience preferences, and seasonal patterns to inform content strategy decisions. (Python, Pandas, Matplotlib, Seaborn)

Supply Chain Analysis — An in-depth look at supply chain data covering product performance, inventory, suppliers, logistics, and customer behavior, with interactive visualizations. (Python, Pandas, NumPy, Plotly)

PreOwned Cars Price Prediction — A regression project predicting used car prices across US regions using Multi-Linear Regression, Decision Trees, Random Forests, and XGBoost, with K-Means clustering for regional segmentation. A follow-up version, V2.0, adds a manufacturer-depreciation-to-odometer metric to improve prediction accuracy. (Python, Scikit-learn, XGBoost)

Synthetic Data Generator — A configurable tool with a Streamlit interface for generating synthetic datasets, useful for testing pipelines and augmenting training data without relying on real user data. (Python, Streamlit)

Website A/B Testing — A data-driven comparison of website performance under light and dark themes, using statistical tests, correlation analysis, and predictive modeling to identify which theme drives stronger engagement. (Python, Pandas, Seaborn, Scikit-learn)



Last Updated: 2026 by PAVAN YELLATHAKOTA

Pinned Loading

  1. BingeMax-Personalized-Movie-Recommendation-Engine BingeMax-Personalized-Movie-Recommendation-Engine Public

    An AI-powered movie recommender using content-based, collaborative, and cosine similarity models. Built with Streamlit + FastAPI.

    Python

  2. KonnectR_flask_fullstack_app KonnectR_flask_fullstack_app Public

    Developed a full-stack web app simulating R\&D collaboration between students, professors, and recruiters with messaging, posting, and analytics modules.

    HTML 1

  3. atmos-db atmos-db Public

    AtmosDB unifies Cloudflare D1, Vectorize, and R2 into a single SDK. Build global, AI-native backends at the edge with a streamlined developer experience and zero dashboard complexity.

    TypeScript

  4. detoxify-telugu detoxify-telugu Public

    A Fine-Tuned BERT-Based Language Model for Hate Speech Detection in Telugu & Tenglish

    Python

  5. pitch-os pitch-os Public