Skip to content

Solutions Architect · Data Engineer · Anaheim, CA

Justin Jenish

Solutions architect for AI & data platforms. With hands still on the pipeline.

I design end-to-end architectures and defend them in front of the people paying for them: city procurement teams, sales leadership, research stakeholders.

Streaming systems and lakehouse architecture, in production since 2020: Kafka and Spark at 5M daily users, Delta Lake tuned for cost, pipelines stress-tested at 50M events/day.

Hiring for

System design defended in front of stakeholders, from procurement boards to sales leadership.Streaming, lakehouse, and pipeline work: throughput, cost, and correctness.

Shipping to production since 2020

Certified across the data stack. MS in Computer Science from USC.

  • University of Southern CaliforniaMS, Computer Science · Viterbi2024
  • SRM Institute of Science & TechnologyBS, Computer Science (Big Data Analytics)2022

The numbers

Production systems since 2020. Measured.

TSL (client: Petco)

Serial-level data integrity across four systems of record

Designing the canonical model, staged legacy migration and reconciliation, and the KPI lineage that lets every reported number trace back to approved source data.

  • Jira Service Management
  • Jira Assets
  • Fishbowl
  • REST APIs
  • Data Migration
  • Audit & Custody

Forward-deployed engineering inside a national retailer's depot operations

Embedded with the client's engineering and operations stakeholders to design an orchestration architecture, a phased and reversible cutover from the incumbent vendor, and role-based experiences that keep depot staff out of admin screens.

  • Solutions Architecture
  • Systems Integration
  • Enterprise Identity (SSO/Entra)
  • Workflow Design
  • Stakeholder Discovery
Read the case study
98%
contract CPI target
4
systems orchestrated
Serial-level
chain of custody
FDE
embedded delivery model
45–60 days
controlled transition window
98%
contract CPI target

USC, Viterbi School of Engineering

A 50M-events/day multi-agent pipeline on Databricks and Airflow

Orchestrated RAG on Databricks with Airflow DAGs and an MLflow retraining lifecycle, stress-tested at 50M events/day against replayed public clickstream data.

  • Databricks
  • Airflow
  • MLflow
  • RAG
  • Multi-Agent
  • Python

Cutting analyst report turnaround from an hour to five minutes

Designed a system that lets non-technical stakeholders get custom analytical reports directly from a plain-language question, removing the analyst queue as a bottleneck across three business domains.

  • Multi-Agent Systems
  • RAG
  • Databricks
  • Airflow
  • MLflow
Read the case study
50M
events/day sustained in load test
~5 min
end-to-end pipeline run
MLflow
automated retraining + tuning
~1 hr → 5 min
report turnaround
3
business domains served
0
analyst hand-offs required

TSL, GSA-cleared government services & electronics contractor

Engineering to federal control requirements

Leading software engineering for government RFP responses in a GSA-cleared environment, designing against FedRAMP control requirements so the security and evidence posture is built into the system rather than retrofitted.

  • FedRAMP-aligned Design
  • Security & Compliance
  • Government RFP
  • Systems Engineering

Defending architecture in front of a city procurement board

Lead engineer for government RFP responses, designing against FedRAMP controls and delivering live product demos to the City of Los Angeles that turned technical capability into stakeholder value and advanced real procurement opportunities.

  • Solutions Architecture
  • Government RFP
  • Stakeholder Demos
  • Security & Compliance
Read the case study
FedRAMP
control requirements designed against
GSA-cleared
contracting environment
Lead engineer
on RFP responses
City of LA
live stakeholder demos
FedRAMP
control-aligned design
GSA-cleared
contracting environment

Aipower Inc.

Hybrid dense + sparse retrieval that fixed exact-term failure

Replaced a dense-only baseline with hybrid retrieval (semantic search for intent, keyword search for SKUs and model numbers), improving precision and cutting latency simultaneously.

  • RAG
  • Hybrid Retrieval
  • LLMs
  • Vector Search
  • Python

Requirements from the revenue side, translated into an AI solution design

Led technical discovery with marketing and sales, turned their customer requirements into AI solution designs and architecture decisions, and delivered the contracted product ahead of the department wind-down.

  • Solution Design
  • Technical Discovery
  • RAG
  • Retrieval Architecture
Read the case study
Dense + sparse
hybrid retrieval architecture
↓ latency
versus dense-only baseline
↑ precision
on SKU / model-number queries
Delivered early
ahead of department wind-down
Sales + marketing
discovery partners
Discovery-led
requirements to architecture

All work

The same nine projects. Ordered for what you’re hiring for.

Petco Depot Operations Platform

Aug 2026 – Present

TSL (client: Petco)

Serial-level data integrity across four systems of record

Designing the canonical model, staged legacy migration and reconciliation, and the KPI lineage that lets every reported number trace back to approved source data.

98%
contract CPI target
4
systems orchestrated
Serial-level
chain of custody
  • Jira Service Management
  • Jira Assets
  • Fishbowl
  • REST APIs
  • Data Migration

Forward-deployed engineering inside a national retailer's depot operations

Embedded with the client's engineering and operations stakeholders to design an orchestration architecture, a phased and reversible cutover from the incumbent vendor, and role-based experiences that keep depot staff out of admin screens.

FDE
embedded delivery model
45–60 days
controlled transition window
98%
contract CPI target
  • Solutions Architecture
  • Systems Integration
  • Enterprise Identity (SSO/Entra)
  • Workflow Design
  • Stakeholder Discovery
Read the case study

AMAR

Sep 2024 – Jun 2025

USC, Viterbi School of Engineering

A 50M-events/day multi-agent pipeline on Databricks and Airflow

Orchestrated RAG on Databricks with Airflow DAGs and an MLflow retraining lifecycle, stress-tested at 50M events/day against replayed public clickstream data.

50M
events/day sustained in load test
~5 min
end-to-end pipeline run
MLflow
automated retraining + tuning
  • Databricks
  • Airflow
  • MLflow
  • RAG
  • Multi-Agent

Cutting analyst report turnaround from an hour to five minutes

Designed a system that lets non-technical stakeholders get custom analytical reports directly from a plain-language question, removing the analyst queue as a bottleneck across three business domains.

~1 hr → 5 min
report turnaround
3
business domains served
0
analyst hand-offs required
  • Multi-Agent Systems
  • RAG
  • Databricks
  • Airflow
  • MLflow
Read the case study

Hybrid-Retrieval Sales Chatbot

Jun 2025 – Oct 2025

Aipower Inc.

Hybrid dense + sparse retrieval that fixed exact-term failure

Replaced a dense-only baseline with hybrid retrieval (semantic search for intent, keyword search for SKUs and model numbers), improving precision and cutting latency simultaneously.

Dense + sparse
hybrid retrieval architecture
↓ latency
versus dense-only baseline
↑ precision
on SKU / model-number queries
  • RAG
  • Hybrid Retrieval
  • LLMs
  • Vector Search
  • Python

Requirements from the revenue side, translated into an AI solution design

Led technical discovery with marketing and sales, turned their customer requirements into AI solution designs and architecture decisions, and delivered the contracted product ahead of the department wind-down.

Delivered early
ahead of department wind-down
Sales + marketing
discovery partners
Discovery-led
requirements to architecture
  • Solution Design
  • Technical Discovery
  • RAG
  • Retrieval Architecture
Read the case study

Social Platform & Revenue Engine

Oct 2025 – Present

TSL (internal venture)

AI and an auditable revenue engine, built from scratch

Building the AI surface and an advertising revenue-share engine whose settlement records are append-only and advertiser-auditable, so the numbers advertisers pay against can be verified rather than trusted.

Blockchain
advertiser-auditable settlement
Front · Back · AI
full ownership
Founding
engineer
  • Flutter/Dart
  • Backend
  • AI
  • Blockchain
  • Revenue Systems

Founding engineer across every layer of a new platform

Owning the architecture end to end: frontend, backend, AI, and the advertising and revenue-share engine, for a social platform aimed at the 2028 Olympics, keeping the whole system coherent while the product was still taking shape.

3 surfaces
frontend · backend · AI
2028 Olympics
target audience
Revenue engine
designed and owned
  • Solutions Architecture
  • Full-Stack
  • Revenue Systems
  • Flutter/Dart
  • AI
Read the case study

Government RFP Delivery & City of LA Demos

Oct 2025 – Present

TSL, GSA-cleared government services & electronics contractor

Engineering to federal control requirements

Leading software engineering for government RFP responses in a GSA-cleared environment, designing against FedRAMP control requirements so the security and evidence posture is built into the system rather than retrofitted.

FedRAMP
control requirements designed against
GSA-cleared
contracting environment
Lead engineer
on RFP responses
  • FedRAMP-aligned Design
  • Security & Compliance
  • Government RFP
  • Systems Engineering

Defending architecture in front of a city procurement board

Lead engineer for government RFP responses, designing against FedRAMP controls and delivering live product demos to the City of Los Angeles that turned technical capability into stakeholder value and advanced real procurement opportunities.

City of LA
live stakeholder demos
FedRAMP
control-aligned design
GSA-cleared
contracting environment
  • Solutions Architecture
  • Government RFP
  • Stakeholder Demos
  • Security & Compliance
Read the case study

Subscriber Streaming Analytics

Dec 2020 – Jun 2021

Lemonpeak (client: a major US satellite/streaming TV provider)

Kafka and Spark Streaming on a 5M-daily-user platform

Lambda-architecture pipelines (Kafka and Spark Streaming for the speed layer, batch for correctness) on AWS EC2/S3, with data-quality assertions via Great Expectations.

5M
daily users on the platform
Lambda
speed + batch architecture
2
junior engineers mentored
  • Kafka
  • Spark Streaming
  • AWS EC2/S3
  • pytest
  • Great Expectations

One platform serving operations, ad-ROI, and the executive team

Designed a two-layer architecture that reconciled real-time operational signal with the reprocessable history that ad-ROI tracking, A/B testing, and executive dashboards required.

3
consumer groups served
5M
daily users
2
engineers mentored
  • Lambda Architecture
  • Distributed Systems
  • AWS
  • Data Quality
Read the case study

Medallion Clickstream Cost Optimization

Jan 2020 – Aug 2020

Arcadia Solutions Inc.

~20% storage and compute savings from file-layout work

Partitioning, file compaction, and indexing across a Delta Lake storage layer on a Medallion clickstream platform: cost down, workloads untouched.

~20%
storage + compute cost cut
Medallion
clickstream architecture
Delta Lake
storage layer tuned
  • Delta Lake
  • Medallion Architecture
  • Partitioning
  • Compaction
  • Spark

Reducing platform cost without touching the consumers

Diagnosed rising lakehouse cost as a storage-layout problem rather than a capacity problem, and cut roughly 20% while predictive sales models and user-behavior analytics kept running.

~20%
cost reduction
0
consumer-facing changes
Lakehouse
cost governance
  • Lakehouse Architecture
  • Cost Optimization
  • Delta Lake
  • Data Modeling
Read the case study

Annenberg Operations Analytics

May 2023 – May 2024

USC, Annenberg School of Journalism & Communication

Snowflake pipelines and a data-vault model over messy operational sources

Built the pipelines, the data-vault schema, and the automated Power BI and Streamlit reporting layer, plus prescriptive ML for staffing optimization.

Data Vault
schema over unstable sources
Automated
Power BI + Streamlit reporting
↓ downtime
A/V and network
  • Snowflake
  • Data Vault
  • Power BI
  • Streamlit
  • Prescriptive ML

Weekly stakeholder reviews turning operational pain into a data model

Owned the analysis, the architecture, and the stakeholder relationship, leading weekly reviews while reducing overdue-equipment incidents and A/V downtime.

Weekly
stakeholder review cadence
↓ response time
staffing + CRM workflows
↓ incidents
equipment overdue
  • Stakeholder Management
  • Data Modeling
  • Snowflake
  • Server Administration
Read the case study

Spatiotemporal GNN for PM2.5 Forecasting

Jan 2022 – May 2022

SRM Institute of Science & Technology

GINs + GRUs beating temporal baselines on a 184-city graph

A spatiotemporal GNN in PyTorch Geometric that models spatial coupling between cities, improving both error and threshold-detection metrics over baseline.

−7.2%
RMSE versus baseline
+12.4%
CSI versus baseline
184
cities modeled
  • PyTorch Geometric
  • GNN
  • GIN
  • GRU
  • Python

Choosing a graph model because the problem was actually spatial

Recognized that temporal baselines structurally could not capture inter-city coupling, and selected an architecture matched to the problem's shape, validated on a four-day horizon.

−7.2%
RMSE
4-day
prediction horizon
184
cities
  • Model Architecture
  • Research
  • PyTorch
  • Graph Modeling
Read the case study

Background

Production since 2020. USC in the middle.

Experience

  • Lead Full-Stack Engineer & Solutions Architect

    Oct 2025 – Present

    TSL, GSA-cleared government services & electronics contractor · Los Angeles, CA

    • Forward-deployed engineer on TSL's IT service and repair contract with Petco, leading architecture and technical discovery for a depot operations platform: an integration and workflow layer orchestrating Jira Service Management, Jira Assets, Fishbowl Inventory, and carrier APIs behind role-based interfaces, with serial-level chain of custody and three contract CPIs targeted at 98%.
    • Lead software engineering for government RFP responses, designing against FedRAMP control requirements; delivered live product demos to the City of Los Angeles, translating technical capability into stakeholder-facing value and advancing procurement opportunities.
    • Founding engineer of an internal social-platform venture, owning architecture and delivery across frontend (Flutter/Dart), backend, and AI; designed the advertising and revenue-share engine, using blockchain-based settlement records for advertiser-auditable spend.
    • Sole software engineer on an autonomous mobile battery-unit prototype, building control software and systems integration for field-deployable energy delivery.
    • Leading SOC 2 readiness across two SaaS products, working through the control and evidence requirements ahead of audit.
  • AI Engineer

    Jun 2025 – Oct 2025

    Aipower Inc., AI products startup · Los Angeles, CA

    • Improved chatbot answer precision on exact-term queries (SKUs, model numbers) and cut response latency versus a dense-only baseline by architecting hybrid retrieval for AI sales chatbots in e-commerce guided selling and troubleshooting.
    • Led technical discovery with marketing and sales teams, translating customer requirements into AI solution designs and architecture decisions; delivered the contracted product ahead of the department wind-down.
    • Built a full-stack, event-driven workflow-management tool coordinating customer service, sales, and inventory across departments.
  • Data Engineer (Research)

    Sep 2024 – Jun 2025

    USC, Viterbi School of Engineering · Los Angeles, CA

    • Built AMAR, a multi-agent RAG pipeline that turns plain-language questions from non-technical stakeholders into custom analytical reports, with automated retraining and parameter tuning through the Databricks MLflow model lifecycle.
    • Cut report turnaround from ~1 hour to 5 minutes across financial-reporting, healthcare-diagnostics, and customer-support scenarios by orchestrating the pipeline on Databricks and Airflow DAGs, stress-tested at 50M events/day with replayed public clickstream data.
  • Business Analyst / Tech Consultant

    May 2023 – May 2024

    USC, Annenberg School of Journalism & Communication · Los Angeles, CA

    • Reduced equipment-overdue incidents and A/V and network downtime by building Snowflake pipelines, a data-vault schema, and automated Power BI/Streamlit dashboards; cut response times by streamlining staffing and CRM workflows.
    • Developed prescriptive ML models for staffing optimization; led weekly stakeholder reviews and managed server administration with a focus on security.
Earlier work: 5 concurrent roles during undergraduate study (2020–2022)
  • Freelance Data Analyst (Finance)

    Jul 2021 – Jul 2022

    GMG Associates · Bangalore, IN

    • Eliminated 10+ hours/day of manual reporting across the team by automating ERP, CRM, and tax-data migration to Azure with Airflow, validation scripts, and unit tests; implemented role-based access for compliance and audit readiness; built Power BI reports and predictive models on Databricks (MLlib, MLflow).
  • Deep Learning Research Assistant

    Jan 2022 – May 2022

    SRM Institute of Science & Technology · Chennai, IN

    • Improved PM2.5 forecasting accuracy (7.2% lower RMSE, 12.4% higher CSI over baseline models, with 4-day prediction horizons) by engineering a spatiotemporal GNN architecture (GINs + GRUs, PyTorch Geometric) across 184 cities.
  • Data Engineer (Streaming Data)

    Dec 2020 – Jun 2021

    Lemonpeak, staffing firm; client: a major US satellite/streaming TV provider · Chennai, IN

    • Built Lambda-architecture pipelines (Kafka, Spark Streaming, AWS EC2/S3) for subscriber analytics on the client's 5M-daily-user platform; supported ad-ROI tracking, A/B testing, and executive dashboards; mentored 2 junior engineers on testing and warehousing practices (pytest, Great Expectations).
  • Software Development Intern

    Feb 2021 – Jun 2021

    Yatnam Technologies · Kochi, IN

    • Built a UK e-commerce platform (warehouse management + ERP) on REST APIs and ETL pipelines; deployed with Terraform and CI/CD on Azure.
  • Data Engineer

    Jan 2020 – Aug 2020

    Arcadia Solutions Inc. · Remote

    • Cut storage and compute costs ~20% on a Medallion clickstream platform by tuning partitioning, file compaction, and indexing across the Delta Lake storage layer, while supporting predictive sales models and user-behavior analytics.

Education

  • MS, Computer Science

    University of Southern California

    Viterbi School of Engineering

    Los Angeles, CA · May 2024

  • BS, Computer Science (Big Data Analytics)

    SRM Institute of Science & Technology

    Chennai, IN · May 2022

Certifications

Technical skills

Data Engineering
SparkPySparkKafkaAirflowFlinkdbtAWS LambdaGlueETL/ELTStreamingREST APIs
Storage & Warehousing
SnowflakeDatabricksDelta LakeRedshiftPostgreSQLS3
Cloud & Architecture
AzureAWSGCPMicrosoft FabricLakehouse & MedallionData ModelingDistributed SystemsSystem DesignCloud MigrationCost OptimizationSystems Integration
AI & Machine Learning
LLMsRAGMulti-Agent SystemsPyTorchMLflowMLlibPredictive & Prescriptive Modeling
Languages & Libraries
PythonSQLScalaJavaScriptTypeScriptDartBashPandasNumPy
Infrastructure & DevOps
LinuxGitDockerTerraformKubernetes (EKS)HelmPrometheusCI/CD
Data Governance & Quality
Unity CatalogCollibraGreat ExpectationsIAM & Role-Based AccessCompliance & Audit ReadinessSOC 2FedRAMP
Visualization & BI
Power BIStreamlit
Collaboration
Requirements GatheringStakeholder ManagementTechnical LeadershipAgile / JIRA

About

I got here by taking things apart. A few even went back together.

It started with toys, moved on to drones, and had turned into robotics by the time I started my undergraduate degree. There was never a career plan behind any of it. I just wanted to know how the thing worked, and the fastest way to find out was usually to open it.

Most of what I built in college aimed at problems I could see from where I was standing: low-cost robotic prosthetics, traffic signals that could clear a path for an ambulance, a spatiotemporal model for forecasting air quality. Covid and a complete absence of funding ended nearly all of it. What stayed with me was the pattern underneath. Every one of those projects lived or died on data: data to find the problem, data to study it, data to know whether the fix actually worked.

Which is a slightly unromantic conclusion for someone who wanted to build robots. It is also what sent me to USC for a Master's, where I specialized in data at scale and in the less glamorous half of it: governance, and using it carefully.

The career took the scenic route from there. Business analyst, data engineer, AI engineer, now lead engineer and architect, across finance, e-commerce, streaming, energy, and social. None of it was planned. But it left me with a suspicion I keep finding evidence for: most hard engineering problems turn out to be translation problems between people who do not share a vocabulary. A good part of my job is standing in that gap with a whiteboard.

These days the work splits in two directions. One is a social platform being built for the 2028 Olympics. The other is a GSA-cleared government contractor whose project list reads a little like a dare: datacenter buildouts, off-grid power systems, robotics, graphene-based materials, federal RFPs, and, inevitably, software.

Alongside that I run my own research in applied ML for energy systems: demand flexibility for AI datacenters, battery life prediction that transfers across cell chemistries, and forecasting that stays calibrated when the distribution shifts underneath it. With a friend at UCLA I also work on AI safety and ethics. Both are being written toward peer-reviewed publication. Neither is there yet, and getting there is proving to be a slower and more humbling education than the research itself.

When I am not doing that, I am usually building RC cars or outside doing something. Ideally both at once.

Justin Jenish