Skip to content
DataCareerHub.ioLearn · Prepare · Apply
Technical Lead

Dennis Mathew Jose

Lead Data Scientist, DataReady AI Lab

Dennis connects production data engineering, machine learning, dashboards, and GenAI/RAG planning into practical AI-readiness work for organizations that need reliable analytics before investing in AI tools.

LLM/RAGProduction pipelines
SnowflakeETL and data modeling
TableauSelf-serve dashboards
DataReady AI Lab Role

Practical technical leadership for AI-ready data audits.

Dennis leads the technical review behind data quality scoring, dashboard-readiness planning, AI/RAG-readiness scoring, and implementation planning. He is not listed as founder.

DataReady Responsibilities

These are the areas Dennis owns technically for the audit service and client-readiness work.

Data quality assessment
Python/SQL analysis
Dashboard planning
AI-readiness scoring
RAG/document-search planning
Technical recommendations
Case study development

Why This Background Matters

Dennis' work spans production LLM pipelines, Snowflake ETL, Tableau dashboards, statistical anomaly detection, clinical deep learning research, and backend financial systems. For DataReady AI Lab, that background supports cautious, practical audits that focus on data quality, KPI clarity, dashboard readiness, automation opportunities, and AI/document-search planning.

The goal is not to promise a finished AI assistant during an audit. The goal is to identify whether the organization's data and documents are clean, structured, governed, and suitable for dashboards, automation, or private AI/RAG planning.

Experience

Professional Experience

Sep 2025 - Apr 2026

Data Scientist (Student Admin Fellow)

Northeastern University, College of Professional Studies
  • Deployed production LLM pipelines using Anthropic APIs with prompt engineering and evaluation cycles for reliable institutional reporting outputs.
  • Built automated Snowflake data lake ETL pipelines and self-serve Tableau dashboards with layered data modeling, cutting reporting turnaround by 93%.
  • Embedded statistical anomaly detection, data quality validation, and feature engineering into pipelines for governance compliance and schema consistency.
  • Translated ambiguous stakeholder needs into structured GenAI and predictive analytics solutions through consultative, hypothesis-driven analysis.
Sep 2025 - Dec 2025

Machine Learning Research (Experiential Learning)

Zeta Surgical x Northeastern University, Boston, MA
  • Designed and evaluated ResNet, DenseNet, and Vision Transformer models for clinical image classification across seven diagnostic categories.
  • Architected a research pipeline from data lake ingestion through feature engineering, model training, and evaluation with reproducibility as a core constraint.
  • Resolved extreme class imbalance with Dice+BCE combined loss and progressive fine-tuning, documenting tradeoffs with ablation studies.
Jul 2023 - Jul 2024

Associate Consultant

Oracle Financial Services, Bengaluru, India
  • Developed backend systems using Java, Python, and PL/SQL for high-volume financial transaction processing, improving system throughput by 30%.
  • Built automated data processing pipelines and batch-processing logic that reduced query execution time by 20% during peak loads.
  • Collaborated with engineering and QA teams to debug production code, improve performance, and maintain system reliability.
Skills

Technical Skills

ML Frameworks & GenAI

PyTorch Hugging Face LLMs and fine-tuning RAG Prompt engineering scikit-learn

Deep Learning

Neural networks Transformers Vision Transformers Attention mechanisms Bayesian methods

Model Optimization

MC Dropout Deep ensembles Uncertainty quantification Inference efficiency

Software Engineering

Python Docker GitHub CI/CD Pytest Git Automated testing GCP AWS

Data & Pipelines

SQL Snowflake NoSQL Java C++ Data preprocessing Feature engineering
Research & Engineering

Selected Projects

May 2026

Uncertainty-Aware Molecular Property Prediction

PyTorch, ChemBERTa, RDKit, XGBoost, MC Dropout, GCP

Designed a multi-modal inference pipeline integrating ChemBERTa transformer embeddings, Morgan fingerprints, and gradient-boosted models, with uncertainty methods to flag low-confidence predictions.

Dec 2025

CiteConnect: Production GenAI System with RAG

LLaMA 3, LangChain, PgVector, Python, Docker, CI/CD

Architected and deployed a hybrid RAG system with semantic search, prompt engineering, containerized deployment, automated testing, and retrieval evaluation using offline and online metrics.