Skip to content

Solutions Architect · Data Engineer · Anaheim, CA

Justin Jenish

Solutions architect for AI & data platforms. With hands still on the pipeline.

I design data and AI platforms end to end and stay hands-on building them: streaming and lakehouse systems in production since 2020, and the architectures I defend in front of the people paying for them, from city procurement boards to sales leadership.

Shipping to production since 2020

Certified across the data stack. MS in Computer Science from USC.

  • University of Southern CaliforniaMS, Computer Science · Viterbi2024
  • SRM Institute of Science & TechnologyBS, Computer Science (Big Data Analytics)2022

The numbers

Production systems since 2020. Measured.

TSL (client: Petco)

One orchestration layer over four systems that each own their own truth

As forward-deployed engineer I'm designing the integration layer for Petco's device depot, mid-transition off its incumbent vendor. JSM, Jira Assets and Fishbowl each stay authoritative over what they own; the binding constraint is serial-level chain of custody, so the migration quarantines unknown records rather than guessing. A wrong serial that looks right is worse than one flagged for reconciliation.

  • Solutions Architecture
  • Systems Integration
  • Jira Service Management
  • Jira Assets
  • Fishbowl
  • Data Migration
Read the case study
4
systems of record orchestrated
Serial-level
chain of custody
98%
contract CPI target

USC, Viterbi School of Engineering

From a plain-language question to a finished analytical report in five minutes

A multi-agent RAG pipeline on Databricks and Airflow that parses a question, plans the analysis, and writes the report end to end, cutting turnaround from roughly an hour of analyst work to about five minutes across three domains. Retraining runs through the MLflow lifecycle rather than manual redeploys, and I load-tested it at 50M events/day on replayed clickstream before trusting the number.

  • Databricks
  • Airflow
  • MLflow
  • RAG
  • Multi-Agent Systems
  • Python
Read the case study
~1 hr → 5 min
report turnaround
50M
events/day in load test
3
business domains served

TSL, GSA-cleared government services & electronics contractor

Government RFPs designed to federal controls and defended in live demos

Leading engineering for government RFP responses in a GSA-cleared environment, designed against FedRAMP control requirements from the start so security is a property of the architecture rather than a retrofit. I deliver the City of Los Angeles demos myself; in procurement, a working demo collapses doubt a document cannot. The RFPs remain in evaluation, so this is control-aligned design, not a completed authorization.

  • Solutions Architecture
  • Government RFP
  • Security & Compliance
  • Stakeholder Demos
Read the case study
City of LA
live stakeholder demos
FedRAMP
control-aligned design
GSA-cleared
contracting environment

Aipower Inc.

Hybrid retrieval that stopped an e-commerce chatbot from fumbling SKUs

Dense-only semantic search failed exactly where e-commerce answers matter most: exact terms like SKUs and model numbers, where embedding similarity blurs the tokens that distinguish products. Adding a sparse keyword path alongside the dense one fixed the failure at its source and cut latency versus the dense-only baseline. The requirements came from discovery with sales and marketing, who knew which queries were losing customers.

  • RAG
  • Hybrid Retrieval
  • LLMs
  • Vector Search
  • Python
Read the case study
Dense + sparse
hybrid retrieval
↓ latency
vs dense-only baseline
↑ precision
on SKU / model-number queries

All work

Nine projects. Strongest first.

Petco Depot Operations Platform

Aug 2026 – Present

TSL (client: Petco)

One orchestration layer over four systems that each own their own truth

As forward-deployed engineer I'm designing the integration layer for Petco's device depot, mid-transition off its incumbent vendor. JSM, Jira Assets and Fishbowl each stay authoritative over what they own; the binding constraint is serial-level chain of custody, so the migration quarantines unknown records rather than guessing. A wrong serial that looks right is worse than one flagged for reconciliation.

4
systems of record orchestrated
Serial-level
chain of custody
98%
contract CPI target
  • Solutions Architecture
  • Systems Integration
  • Jira Service Management
  • Jira Assets
  • Fishbowl
Read the case study

Government RFP Delivery & City of LA Demos

Oct 2025 – Present

TSL, GSA-cleared government services & electronics contractor

Government RFPs designed to federal controls and defended in live demos

Leading engineering for government RFP responses in a GSA-cleared environment, designed against FedRAMP control requirements from the start so security is a property of the architecture rather than a retrofit. I deliver the City of Los Angeles demos myself; in procurement, a working demo collapses doubt a document cannot. The RFPs remain in evaluation, so this is control-aligned design, not a completed authorization.

City of LA
live stakeholder demos
FedRAMP
control-aligned design
GSA-cleared
contracting environment
  • Solutions Architecture
  • Government RFP
  • Security & Compliance
  • Stakeholder Demos
Read the case study

AMAR

Sep 2024 – Jun 2025

USC, Viterbi School of Engineering

From a plain-language question to a finished analytical report in five minutes

A multi-agent RAG pipeline on Databricks and Airflow that parses a question, plans the analysis, and writes the report end to end, cutting turnaround from roughly an hour of analyst work to about five minutes across three domains. Retraining runs through the MLflow lifecycle rather than manual redeploys, and I load-tested it at 50M events/day on replayed clickstream before trusting the number.

~1 hr → 5 min
report turnaround
50M
events/day in load test
3
business domains served
  • Databricks
  • Airflow
  • MLflow
  • RAG
  • Multi-Agent Systems
Read the case study

Subscriber Streaming Analytics

Dec 2020 – Jun 2021

Lemonpeak (client: a major US satellite/streaming TV provider)

Lambda-architecture analytics for a 5M-user streaming platform

Subscriber analytics for a satellite and streaming TV provider that had to serve two incompatible needs at once: real-time signal for operations, and correct, reprocessable history for ad-ROI and executive reporting. I built Lambda-architecture pipelines on Kafka and Spark Streaming, a speed layer for latency and a batch layer for correctness, with data-quality assertions in the pipeline because at 5M daily users broken data corrupts a dashboard more quietly than broken code.

5M
daily users on the platform
Lambda
speed + batch architecture
2
engineers mentored
  • Kafka
  • Spark Streaming
  • AWS EC2/S3
  • Great Expectations
  • pytest
Read the case study

Hybrid-Retrieval Sales Chatbot

Jun 2025 – Oct 2025

Aipower Inc.

Hybrid retrieval that stopped an e-commerce chatbot from fumbling SKUs

Dense-only semantic search failed exactly where e-commerce answers matter most: exact terms like SKUs and model numbers, where embedding similarity blurs the tokens that distinguish products. Adding a sparse keyword path alongside the dense one fixed the failure at its source and cut latency versus the dense-only baseline. The requirements came from discovery with sales and marketing, who knew which queries were losing customers.

Dense + sparse
hybrid retrieval
↓ latency
vs dense-only baseline
↑ precision
on SKU / model-number queries
  • RAG
  • Hybrid Retrieval
  • LLMs
  • Vector Search
  • Python
Read the case study

Social Platform & Revenue Engine

Oct 2025 – Present

TSL (internal venture)

Founding engineer on a social platform, from the frontend to an auditable ad economy

Building a social platform aimed at the 2028 Olympics as founding engineer, owning frontend (Flutter/Dart), backend, and AI. The piece I care most about is the revenue engine: advertisers will not trust spend the platform counts itself, so settlement runs on append-only, advertiser-auditable records where the split is verifiable by the party with the most reason to doubt it.

Front · Back · AI
full ownership
Blockchain
advertiser-auditable settlement
2028 Olympics
target audience
  • Flutter/Dart
  • Backend
  • AI
  • Blockchain
  • Revenue Systems
Read the case study

Medallion Clickstream Cost Optimization

Jan 2020 – Aug 2020

Arcadia Solutions Inc.

Cutting lakehouse cost ~20% by fixing file layout, not buying compute

A Medallion clickstream platform was accumulating cost faster than value, the classic small-files and poor-partitioning tax on a Delta Lake. Rising lakehouse cost usually looks like a compute problem and is actually a file-layout one, so I tuned partitioning to real query predicates, compacted small files, and added indexing, cutting storage and compute roughly 20% while the predictive-sales and user-behavior workloads on top kept running untouched.

~20%
storage + compute cost cut
0
consumer-facing changes
Delta Lake
storage layer tuned
  • Delta Lake
  • Medallion Architecture
  • Spark
  • Partitioning
  • Compaction
Read the case study

Annenberg Operations Analytics

May 2023 – May 2024

USC, Annenberg School of Journalism & Communication

A data-vault model that tamed messy A/V and equipment operations

Equipment was going overdue and A/V and network downtime was going unmanaged because the operational data lived in disconnected systems with no shared model. I built Snowflake pipelines and a data-vault schema, chosen over a direct star schema because the sources changed shape often and data vault absorbs that without rewriting history, then automated Power BI and Streamlit reporting on top and led the weekly stakeholder reviews that turned the pain into requirements.

Data Vault
schema over unstable sources
↓ downtime
A/V and network
Weekly
stakeholder reviews led
  • Snowflake
  • Data Vault
  • Power BI
  • Streamlit
  • Prescriptive ML
Read the case study

Spatiotemporal GNN for PM2.5 Forecasting

Jan 2022 – May 2022

SRM Institute of Science & Technology

A spatiotemporal GNN that forecasts air quality across 184 cities

PM2.5 is spatially coupled (a city's air quality depends on its neighbors') which purely temporal baselines structurally cannot represent. I modeled 184 cities as a graph and combined Graph Isomorphism Networks with GRUs in PyTorch Geometric, matching the architecture to the shape of the problem, and beat the baselines by 7.2% RMSE and 12.4% CSI at four-day horizons.

−7.2%
RMSE vs baseline
+12.4%
CSI vs baseline
184
cities modeled
  • PyTorch Geometric
  • GNN
  • GIN
  • GRU
  • Python
Read the case study

Background

Production since 2020. USC in the middle.

Experience

  • Lead Full-Stack Engineer & Solutions Architect

    Oct 2025 – Present

    TSL, GSA-cleared government services & electronics contractor · Los Angeles, CA

    • Forward-deployed engineer on TSL's IT service and repair contract with Petco, leading architecture and technical discovery for a depot operations platform: an integration and workflow layer orchestrating Jira Service Management, Jira Assets, Fishbowl Inventory, and carrier APIs behind role-based interfaces, with serial-level chain of custody and three contract CPIs targeted at 98%.
    • Lead software engineering for government RFP responses, designing against FedRAMP control requirements; delivered live product demos to the City of Los Angeles, translating technical capability into stakeholder-facing value and advancing procurement opportunities.
    • Founding engineer of an internal social-platform venture, owning architecture and delivery across frontend (Flutter/Dart), backend, and AI; designed the advertising and revenue-share engine, using blockchain-based settlement records for advertiser-auditable spend.
    • Sole software engineer on an autonomous mobile battery-unit prototype, building control software and systems integration for field-deployable energy delivery.
    • Leading SOC 2 readiness across two SaaS products, working through the control and evidence requirements ahead of audit.
  • AI Engineer

    Jun 2025 – Oct 2025

    Aipower Inc., AI products startup · Los Angeles, CA

    • Improved chatbot answer precision on exact-term queries (SKUs, model numbers) and cut response latency versus a dense-only baseline by architecting hybrid retrieval for AI sales chatbots in e-commerce guided selling and troubleshooting.
    • Led technical discovery with marketing and sales teams, translating customer requirements into AI solution designs and architecture decisions; delivered the contracted product ahead of the department wind-down.
    • Built a full-stack, event-driven workflow-management tool coordinating customer service, sales, and inventory across departments.
  • Data Engineer (Research)

    Sep 2024 – Jun 2025

    USC, Viterbi School of Engineering · Los Angeles, CA

    • Built AMAR, a multi-agent RAG pipeline that turns plain-language questions from non-technical stakeholders into custom analytical reports, with automated retraining and parameter tuning through the Databricks MLflow model lifecycle.
    • Cut report turnaround from ~1 hour to 5 minutes across financial-reporting, healthcare-diagnostics, and customer-support scenarios by orchestrating the pipeline on Databricks and Airflow DAGs, stress-tested at 50M events/day with replayed public clickstream data.
  • Business Analyst / Tech Consultant

    May 2023 – May 2024

    USC, Annenberg School of Journalism & Communication · Los Angeles, CA

    • Reduced equipment-overdue incidents and A/V and network downtime by building Snowflake pipelines, a data-vault schema, and automated Power BI/Streamlit dashboards; cut response times by streamlining staffing and CRM workflows.
    • Developed prescriptive ML models for staffing optimization; led weekly stakeholder reviews and managed server administration with a focus on security.
Earlier work: 5 concurrent roles during undergraduate study (2020–2022)
  • Freelance Data Analyst (Finance)

    Jul 2021 – Jul 2022

    GMG Associates · Bangalore, IN

    • Eliminated 10+ hours/day of manual reporting across the team by automating ERP, CRM, and tax-data migration to Azure with Airflow, validation scripts, and unit tests; implemented role-based access for compliance and audit readiness; built Power BI reports and predictive models on Databricks (MLlib, MLflow).
  • Deep Learning Research Assistant

    Jan 2022 – May 2022

    SRM Institute of Science & Technology · Chennai, IN

    • Improved PM2.5 forecasting accuracy (7.2% lower RMSE, 12.4% higher CSI over baseline models, with 4-day prediction horizons) by engineering a spatiotemporal GNN architecture (GINs + GRUs, PyTorch Geometric) across 184 cities.
  • Data Engineer (Streaming Data)

    Dec 2020 – Jun 2021

    Lemonpeak, staffing firm; client: a major US satellite/streaming TV provider · Chennai, IN

    • Built Lambda-architecture pipelines (Kafka, Spark Streaming, AWS EC2/S3) for subscriber analytics on the client's 5M-daily-user platform; supported ad-ROI tracking, A/B testing, and executive dashboards; mentored 2 junior engineers on testing and warehousing practices (pytest, Great Expectations).
  • Software Development Intern

    Feb 2021 – Jun 2021

    Yatnam Technologies · Kochi, IN

    • Built a UK e-commerce platform (warehouse management + ERP) on REST APIs and ETL pipelines; deployed with Terraform and CI/CD on Azure.
  • Data Engineer

    Jan 2020 – Aug 2020

    Arcadia Solutions Inc. · Remote

    • Cut storage and compute costs ~20% on a Medallion clickstream platform by tuning partitioning, file compaction, and indexing across the Delta Lake storage layer, while supporting predictive sales models and user-behavior analytics.

Education

  • MS, Computer Science

    University of Southern California

    Viterbi School of Engineering

    Los Angeles, CA · May 2024

  • BS, Computer Science (Big Data Analytics)

    SRM Institute of Science & Technology

    Chennai, IN · May 2022

Certifications

Technical skills

Data Engineering
SparkPySparkKafkaAirflowFlinkdbtAWS LambdaGlueETL/ELTStreamingREST APIs
Storage & Warehousing
SnowflakeDatabricksDelta LakeRedshiftPostgreSQLS3
Cloud & Architecture
AzureAWSGCPMicrosoft FabricLakehouse & MedallionData ModelingDistributed SystemsSystem DesignCloud MigrationCost OptimizationSystems Integration
AI & Machine Learning
LLMsRAGMulti-Agent SystemsPyTorchMLflowMLlibPredictive & Prescriptive Modeling
Languages & Libraries
PythonSQLScalaJavaScriptTypeScriptDartBashPandasNumPy
Infrastructure & DevOps
LinuxGitDockerTerraformKubernetes (EKS)HelmPrometheusCI/CD
Data Governance & Quality
Unity CatalogCollibraGreat ExpectationsIAM & Role-Based AccessCompliance & Audit ReadinessSOC 2FedRAMP
Visualization & BI
Power BIStreamlit
Collaboration
Requirements GatheringStakeholder ManagementTechnical LeadershipAgile / JIRA

About

I got here by taking things apart. A few even went back together.

It started with toys, moved on to drones, and had turned into robotics by the time I started my undergraduate degree. There was never a career plan behind any of it. I just wanted to know how the thing worked, and the fastest way to find out was usually to open it.

Most of what I built in college aimed at problems I could see from where I was standing: low-cost robotic prosthetics, traffic signals that could clear a path for an ambulance, a spatiotemporal model for forecasting air quality. Covid and a complete absence of funding ended nearly all of it. What stayed with me was the pattern underneath. Every one of those projects lived or died on data: data to find the problem, data to study it, data to know whether the fix actually worked.

Which is a slightly unromantic conclusion for someone who wanted to build robots. It is also what sent me to USC for a Master's, where I specialized in data at scale and in the less glamorous half of it: governance, and using it carefully.

The career took the scenic route from there. Business analyst, data engineer, AI engineer, now lead engineer and architect, across finance, e-commerce, streaming, energy, and social. None of it was planned. But it left me with a suspicion I keep finding evidence for: most hard engineering problems turn out to be translation problems between people who do not share a vocabulary. A good part of my job is standing in that gap with a whiteboard.

These days the work splits in two directions. One is a social platform being built for the 2028 Olympics. The other is a GSA-cleared government contractor whose project list reads a little like a dare: datacenter buildouts, off-grid power systems, robotics, graphene-based materials, federal RFPs, and, inevitably, software.

Alongside that I run my own research in applied ML for energy systems: demand flexibility for AI datacenters, battery life prediction that transfers across cell chemistries, and forecasting that stays calibrated when the distribution shifts underneath it. With a friend at UCLA I also work on AI safety and ethics. Both are being written toward peer-reviewed publication. Neither is there yet, and getting there is proving to be a slower and more humbling education than the research itself.

When I am not doing that, I am usually building RC cars or outside doing something. Ideally both at once.

Justin Jenish