Maharjan Consulting
All case studies
Scientific Computing
Academic research · HPC

Infrastructure dashboards for a 500TB+ research computing environment

Gave researchers real-time visibility into storage and Slurm job queues across a large shared HPC environment.

500TB+
Storage monitored
4
Storage tiers
Slurm
Scheduler

The problem

Researchers sharing a large HPC environment had no clear view of storage consumption across scratch, compute, permanent, and archival tiers, or of where their jobs sat in priority queues. Capacity issues surfaced as failures, not warnings.

The solution

I designed and built internal dashboards that monitor 500TB+ of storage across tiers and visualize Slurm job status by priority queue, so researchers can track usage and manage large-scale analysis workloads proactively.

Technical approach

  • Instrumented storage tiers to collect consumption metrics continuously
  • Integrated Slurm scheduler data into a live job-visualization view
  • Designed dashboards around the questions researchers actually ask
  • Automated data collection so the views stay current without manual work

Business impact

  • Storage pressure became visible before it caused job failures
  • Researchers could see and plan around queue priority
  • Infrastructure decisions moved from anecdote to data
  • Support requests for 'where is my job' dropped
Technology used
Python
Slurm
Linux
HPC
Google Cloud

Have a similar problem?

If this looks like something you're facing, let's talk. Every engagement starts with a free consultation.