DataOps: Data Engineering, Analysis & Science
A solar farm operator's inverter-telemetry scripts, rebuilt as one orchestrated, idempotent pipeline.
Fictionalised summary. Company names, the contact email and repository addresses are placeholders; no credential or token values are shown.
Domain & SKU
What this example orders, at a glance.
- Domain
- DataOps: Data Engineering, Analysis & Science
- Deliverable
- Custom ETL & Data Pipeline
- Turnaround
- Priority Overnight (under 14 hours)
- Environments in scope
- Linux VM
- Compliance
- None apply
02
Current State
Baseline architecture, stack, repository context and known friction.
Architecture & system context
Three scripts run in sequence from a cron job at 2am against our inverter-telemetry CSV drops and a small warehouse database. extract.py loads yesterday's CSV into a staging table, transform.py aggregates it into 15-minute buckets, and load.py upserts into warehouse.db. A shell wrapper chains them together.
Tech stack, frameworks & versions
- Python 3.11
- SQLite (warehouse.db)
- bash cron wrapper
- CSV inverter-telemetry drops
Known issues, error logs & friction
There is no orchestration, logging or retry. If any stage fails the wrapper moves on or stops silently, and we only notice at standup when yesterday's numbers are missing. Re-running by hand can duplicate rows.
Primary repository or architecture link
https://github.com/your-org/your-repoData sources & systems
Daily inverter-telemetry CSV files land in a hardcoded directory on the Linux VM, and the warehouse is a small SQLite database. A sample CSV of about thirty rows is committed in the repository for testing.
03
Target State
Deliverable expectations, measurable benchmarks and definition of done.
Deliverable expectations
One orchestrated pipeline with a single entry point and ordered stages that records stage-level status, writes a structured run log with rows in and out plus duration, marks a run failed with the failing stage named, and makes the load stage idempotent so re-running yesterday is safe.
Quantifiable benchmarks & success metrics
One command runs all stages in order; every run produces a structured log; a failure names the stage; re-running the load stage produces no duplicates.
Definition of done
The orchestrator runs extract, transform and load in order, writes the run log, records failures with the stage, and the idempotency test passes. PIPELINE.md documents how to run and backfill.
Data outcome & insights sought
A reliable nightly pipeline that turns daily inverter CSV drops into 15-minute aggregates in the warehouse, with a run history so the team can see what happened and safely backfill a missed day.
04
Constraints
Forbidden changes, compliance, regions, freeze windows and deadlines.
Forbidden modifications & boundaries
Do not change the warehouse schema beyond what the existing upsert needs, and do not delete or rewrite historical rows.
Compliance standards
- None apply
Approved regions & environments
- Linux VM in the us-west region
Freeze windows & blackout periods
No production changes Fri 17:00 to Mon 06:00 Pacific; the 2am window must remain available.
Pinned libraries, versions & standards
Stay on Python 3.11 and SQLite; no new orchestration framework, keep it simple.
Target deadline or milestone
16 October, before the month-end generation report.
Data governance & privacy constraints
Operational telemetry only, no personal data; keep the raw CSV and warehouse data on the internal VM.
05
Access & Verification
Repository access, read-only credentials, environments and verification.
Handover method
Repository access
Git repository URL
https://github.com/your-org/your-repoRead-only repository token
Encrypted in your browser before it leaves the device.
Environments in scope
- Linux VM
Data sample or schema
data/sample_inverter.csv in the repository (about thirty rows) plus the staging and warehouse schema in the scripts.
Contact email
Acceptance criteria as captured
These are the testable statements fixed before payment. Delivery is checked against this list.
- The three scripts become one orchestrated pipeline: single entry point, ordered stages, stage-level status persisted.
- Every run produces a structured run log (rows in/out per stage, duration) and a failure marks the run failed with the stage named.
- The load stage is idempotent: re-running yesterday is safe and tested.
- `PIPELINE.md` documents operations: how to run, how to backfill, and where logs live.
From the live intake
Captured screenshots of this work order being filled in. Each image is scrollable — scroll inside a frame to see the full page.



What happens after payment
Payment confirms the fixed scope. From there the work runs to the SLA you selected and every step is visible on your private dashboard.
The SLA clock starts
Your turnaround countdown begins the moment payment is confirmed — under 14 hours overnight, or 48 hours standard. Early delivery is always the goal.
A private dashboard
Your tracking link opens a private work-order dashboard showing the five delivery stages and every update as the work progresses.
Verified delivery
The deliverable is returned as a verified pull request or documented package, checked against the acceptance criteria you agreed before payment.
Your downloads
The final report, the acceptance-criteria results and every file in the delivery package are available on the dashboard until you close it.