CASE STUDY · HPC · PARALLEL DATA SCIENCE

Parallel pipelines for
large-scale scientific and ML workflows.

My parallel-computing work spans scientific HPC pipelines at Brunel University and the architecture of machine-learning/data workflows designed to move from local experimentation into reproducible, distributed execution.

PythonLinuxSlurmARCHER HPCHadoop conceptsBatch processingHDF5 / SQLite

HPC RESEARCH WORK

Orchestrating large computational jobs rather than running one notebook at a time.

In bioinformatics/HPC research work I supported automated protein-analysis pipelines using Linux, Python orchestration, Bash and Slurm scheduling in distributed environments including ARCHER. The work involved batch-processing operations, workflow monitoring, debugging, reconciliation and optimisation across large biological datasets.

PIPELINE THINKING

Parallelisation changes how the whole workflow is designed.

A parallel pipeline needs explicit job boundaries, deterministic inputs and outputs, resource-aware scheduling, resumability and failure visibility. Those requirements are different from a single-process data-science script and are directly relevant to large AI/ML workloads such as embedding generation, document processing, evaluation and batch inference.

  • Partition work into independently executable units.
  • Use job scheduling instead of manual sequential execution.
  • Persist intermediate outputs so failed jobs can resume.
  • Track logs and job state for debugging at scale.
  • Separate compute-heavy stages from interactive analysis.

ML / DATA ARCHITECTURE

From synthetic data generation to distributed-ready analytics.

My MSc work also explored a modular pipeline where synthetic data generation, EDA, clustering/classification, server-hosted datasets and dashboard delivery were separated into stages. Python/pandas handled integration; ML components included DBSCAN, PCA, SVM and Random Forest; the architecture explored Hadoop-style storage and Ubuntu/Apache hosting for reproducible experimentation.

WHY THIS MATTERS FOR AI

Many AI bottlenecks are orchestration problems.

Modern AI systems often need to process thousands of documents, run evaluation suites, generate embeddings, call multiple models or execute long-running background tasks. The same principles used in HPC—partitioning, scheduling, observability, deterministic artifacts and recovery—make those systems more reliable and cost-aware.

WHAT THIS DEMONSTRATES

AI/ML engineering beyond model selection.

This work demonstrates comfort with the systems side of data science: Linux environments, parallel/batch workflows, orchestration, scientific compute, reproducibility, debugging and the design of pipelines that can grow beyond a local notebook.

Need AI or ML workflows that have to scale beyond one process?

Discuss a systems project →