Dashboards are only as good as the data that feeds them. In many companies the data foundation is fragmented: sources are not connected, transformations happen manually in Excel, pipelines exist only in the heads of individual employees. The result is fragile reporting that breaks with every change to the source data.
I build data infrastructure that supports analytical systems for the long term — ETL/ELT pipelines that are reproducible, documented, and can be operated independently of any single person.
What is data engineering — and when do you need it?
Data engineering refers to building and operating the technical infrastructure that turns raw data from various sources into a form suitable for analysis and reporting.
Data engineering makes sense when:
reporting processes are manual and error-prone
data from multiple systems (CRM, ERP, database, APIs) has to be brought together
existing pipelines are undocumented, untestable or unmaintainable
a new BI system, dashboard or forecasting model needs a clean data foundation
data transformations today depend on individual people who could leave the company
My approach
1. Taking stock
I analyze your existing data landscape: source systems, data flows, manual steps and dependencies. What already exists? What is missing? What should be replaced or simplified?
2. Architecture decision
Depending on data volume, update frequency and budget, I recommend a suitable architecture — from simple SQL views to orchestrated workflows with Apache Airflow. The solution should fit your infrastructure and your team.
3. Development
Building the pipelines with a clear separation of extraction, transformation and loading. All transformation steps are implemented in versioned code — traceable, testable and reproducible. Manual intervention is reduced to a minimum.
4. Documentation and handover
You receive complete documentation of the architecture, the data flows and the individual transformation steps — so your team can operate the infrastructure independently and extend it as needed.
Tools and technologies I use
I work with SQL (PostgreSQL, MS SQL Server), Python and R for data transformation and pipeline development. For orchestration I use Apache Airflow. Docker-based deployments enable portable, reproducible environments. Integration into existing infrastructure — local or cloud-based — is part of the standard approach.
Who I work with
My clients are companies in the EU that want to professionalize their data foundation for analytical purposes — from the first structured pipeline to modernizing grown, fragile data architectures. I work project-based or hourly — remote or on-site in Berlin.
Reference projects and further reading
- Company-Wide Business Intelligence System — complete BI system with a consolidated data foundation across operational, financial and marketing data
- Run Docker Containers Remotely with Airflow — practical article on pipeline orchestration with Apache Airflow
- Using Airflow FileSensor for Triggering ETL Processes — event-driven ETL pipelines with Airflow
Let’s discuss your project
Do you want to rebuild your data pipelines, have existing structures documented, or put a BI project on a solid data foundation?
Pricing & Packages
- Fixed scope, defined in advance
- 1–2 weeks duration
- 1 revision round
- Summary deliverable document
- Post-delivery support
- Iterative collaboration
- Complete project with iterations
- 3–6 weeks duration
- Unlimited revisions within scope
- Full documentation & handover
- 2 weeks post-delivery support (async)
- Ongoing monthly support
- Monthly, min. 3 months
- Continuous iterations
- Living documentation
- Priority access
- €90/h for additional work
- Monthly review call
Hourly rate for ad hoc requests: €90/h. All prices plus VAT.