JobsJump Trading

HPC Data Center Developer

Jump Trading · Chicago, IL · IT Infrastructure + WCW

Posted Aug 8, 2026 · We last checked this listing on Sep 20, 2026

Apply at Jump Trading

Likely interview questions for this role

Written from this job description, not a generic list. Each one notes what the interviewer is really checking.

Behavioral

Tell me about a piece of infrastructure tooling you shipped that you also had to own long term. What broke after launch, and how did you find out?

whether they understand ownership as sustained reliability, not just a launch event

Tell me about a time you used AI tools as part of your actual development workflow, not just for autocomplete. What went into the code that AI wrote, and how did you verify it before it hit production?

genuine daily use of AI tools with a critical eye, not just claiming familiarity

Describe a production incident you root-caused that turned out to be caused by something outside the code itself, like a networking issue or a hardware fault.

depth of Linux and systems troubleshooting skill, and whether root cause analysis is a real habit

This role includes evenings and weekends for coordinated maintenance windows. Tell me about a time you had to be reliably available for something like that, and how you planned around it.

honesty about availability expectations and track record of following through under an on-call type commitment

Technical

Walk me through a hardware onboarding pipeline you built that took a device from racked and cabled to production-ready with little manual work. What did it automate and where did it still need a human?

whether the candidate has actually built end-to-end onboarding automation versus scattered scripts

Tell me about a time you integrated telemetry from a device using IPMI, Redfish, or SNMP. What made that integration harder than it looked at first?

real hands-on experience with hardware management interfaces, not just familiarity with the acronyms

How would you design a capacity planning tool that models current power and cooling utilization and forecasts when a data hall will run out of headroom?

whether they can turn a vague operational need into a concrete data model and tool

Describe how you'd build an outage simulation that shows what happens if a CDU or a cooling loop fails, including how redundancy and failover paths get validated.

depth of understanding of data center power and cooling topology, not just software skill

How do you decide what should be normalized before it lands in a centralized observability platform when you're pulling metrics from several different colo or DC providers with inconsistent data formats?

practical judgment on data pipeline design for heterogeneous telemetry sources

Walk me through how you'd build a Grafana dashboard and alerting rule for an environmental sensor feed, from raw metric ingestion to an alert someone on-call would actually trust.

end-to-end familiarity with Prometheus or InfluxDB style stacks and alerting design, not just dashboard building

If you had to choose between SaltStack, Ansible, and Terraform for managing configuration across thousands of heterogeneous data center devices, how would you decide, and what tradeoffs would you accept?

practical experience choosing and living with infrastructure-as-code tools at scale, not just naming them

How would you design a schema in ClickHouse or MySQL to hold hardware lifecycle and inventory data that gets queried both for day-to-day operations and for longer-term capacity reporting?

database design skill applied to a real operational use case rather than abstract schema knowledge

Situational

Say you need to onboard a new type of rack PDU that doesn't fit your existing discovery and configuration workflow. How would you extend your tooling without breaking support for the devices already onboarded?

ability to design extensible automation rather than one-off fixes

You're two weeks into a project and the HPC Operations lead's requirements have shifted twice. How do you keep building something useful without just chasing the latest ask?

how they manage close, iterative collaboration with non-engineering stakeholders without losing project direction

Practice this interview out loud.

Offer builds a real interview for this exact role at Jump Trading from your resume and this job description, asks the questions one at a time, and tells you what landed. The first one is free.

Practice this out loud

The full job description

As published by Jump Trading.

<p data-renderer-start-pos="2" data-local-id="87e2ab080593">HPC Data Center Production Engineer<br>Location: Chicago or New York (On-site 5 days/week)<br>Jump Trading Group is committed to world class research. We empower exceptional talents in Mathematics, Physics, and Computer Science to seek scientific boundaries, push through them, and apply cutting edge research to global financial markets. Our culture is unique. Constant innovation requires fearlessness, creativity, intellectual honesty, and a relentless competitive streak. We believe in winning together and unlocking unique individual talent by incenting collaboration and mutual respect. At Jump, research outcomes drive more than superior risk adjusted returns. We design, develop, and deploy technologies that change our world, fund start-ups across industries, and partner with leading global research organizations and universities to solve problems.<br>Trading Infrastructure is a global organization of Engineers who architect, build and maintain our world-class infrastructure. From colo design/implementation, to optimizing our exchange connectivity, to building world class low latent Wide Area Networks, we leverage research and automation to consistently adapt and innovate our infrastructure to scale and drive our trading and evolving business.<br>We are looking for an HPC Data Center Production Engineer to build and own the automation and tooling that powers Jump's HPC data center operations. This is a development-heavy role focused on automating the onboarding and lifecycle management of data center hardware—servers, switches, rack PDUs, CDUs, and environmental sensors—and building tools for capacity planning, outage simulation, monitoring, and metrics integration. You will work hand-in-hand with the HPC Planning, Engineering, and Operations leads to turn tooling and monitoring vision into production-ready systems. Heavy, daily use of AI tools is expected to accelerate development and raise the quality bar on everything you build.<br>What You'll Do:<br>Hardware Onboarding Automation</p> <ul class="ak-ul" data-local-id="6a6501b983bc" data-indent-level="1"> <li> <p data-renderer-start-pos="2005" data-local-id="8e62b7bc5954">Design, develop, and maintain automation to onboard new hardware devices into Jump's HPC data centers, including servers, network switches, rack PDUs, CDUs, and environmental sensors.</p> </li> <li> <p data-renderer-start-pos="2192" data-local-id="140dda2f4f46">Build end-to-end provisioning workflows that take hardware from racked-and-cabled through discovery, configuration, validation, and production-ready state with minimal manual intervention.</p> </li> <li> <p data-renderer-start-pos="2384" data-local-id="3c781b8eb196">Extend and adapt onboarding automation as new hardware platforms and device types are introduced.<br>Data Center Tooling Development</p> </li> <li> <p data-renderer-start-pos="2517" data-local-id="bb7a25ce89e9">Develop tools for power and cooling capacity planning—enabling the operations and planning teams to model current utilization, forecast growth, and identify constraints before they become problems.</p> </li> <li> <p data-renderer-start-pos="2718" data-local-id="a5ee657c6754">Build outage simulation tooling to model the impact of power, cooling, or network failures across HPC facilities and validate redundancy/failover configurations.</p> </li> <li> <p data-renderer-start-pos="2883" data-local-id="7db370f868d6">Develop and maintain operational tooling that supports day-to-day data center workflows such as hardware lifecycle tracking, data center inventory/spares, change management, and diagnostics.<br>Monitoring &amp; Metrics Integration</p> </li> <li> <p data-renderer-start-pos="3110" data-local-id="5fcc27172668">Build and maintain monitoring integrations for HPC data center infrastructure—pulling telemetry from servers, switches, PDUs, CDUs, environmental sensors, and facility systems into centralized observability platforms.</p> </li> <li> <p data-renderer-start-pos="3331" data-local-id="b11e595906cd">Integrate metrics feeds from colocation and data center providers into Jump's monitoring stack, normalizing data for alerting and capacity reporting.</p> </li> <li> <p data-renderer-start-pos="3484" data-local-id="3fbb1d0d94ba">Work with the Operations Lead to implement the monitoring and alerting strategy, translating requirements into deployed, production-grade instrumentation.<br>Cross-Team Collaboration</p> </li> <li> <p data-renderer-start-pos="3667" data-local-id="ccfd5d3e9513">Work very closely with the HPC Planning, Engineering, and Operations leads to understand tooling and monitoring needs and bring their vision to fruition.</p> </li> <li> <p data-renderer-start-pos="3824" data-local-id="3998dc058560">Partner with HPC Engineering on integration points between data center automation and compute/storage/network provisioning systems.</p> </li> <li> <p data-renderer-start-pos="3959" data-local-id="e41f4b0272ac">Translate operational pain points and manual processes into automated, maintainable solutions.<br>Systems Maintenance &amp; Reliability</p> </li> <li> <p data-renderer-start-pos="4091" data-local-id="e06e2cba53e8">Own the reliability and lifecycle of all systems and tools you develop—monitor for failures, respond to issues, and iterate based on operational feedback.</p> </li> <li> <p data-renderer-start-pos="4249" data-local-id="4a9205b19f0a">Maintain comprehensive documentation for all tooling, automation workflows, and integrations.</p> </li> <li> <p data-renderer-start-pos="4346" data-local-id="b6e9f13984c1">Participate in large, coordinated maintenance operations, including during evenings and weekends.<br>AI-Driven Development</p> </li> <li> <p data-renderer-start-pos="4469" data-local-id="6ad884397749">Use AI tools daily across all aspects of the role: writing and reviewing code, analyzing data, debugging, generating documentation, and accelerating development velocity.</p> </li> <li> <p data-renderer-start-pos="4643" data-local-id="46d0eb7a8a74">Identify opportunities to apply AI to data center operations problems—anomaly detection, predictive capacity planning, intelligent alerting, and beyond.<br>Additional duties as assigned or needed.<br>Skills You'll Need:</p> </li> <li> <p data-renderer-start-pos="4860" data-local-id="b7996ffc06ff">5+ years of professional experience in production engineering, infrastructure automation, or site reliability engineering, preferably in HPC or large-scale data center environments.</p> </li> <li> <p data-renderer-start-pos="5045" data-local-id="37c1cd214eb4">Proven track record of building and shipping production automation and tooling—not just scripts, but maintained, reliable systems.</p> </li> <li> <p data-renderer-start-pos="5179" data-local-id="7c7ae7b72ba1">Experience automating hardware provisioning and lifecycle management (servers, network devices, power/cooling infrastructure).</p> </li> <li> <p data-renderer-start-pos="5309" data-local-id="da807c0e7bf9">Strong understanding of data center infrastructure: power distribution, cooling systems (air and liquid), environmental monitoring, and structured cabling.</p> </li> <li> <p data-renderer-start-pos="5468" data-local-id="8b7681c7c4d2">Experience integrating with hardware management interfaces (IPMI/BMC/Redfish, SNMP, vendor APIs) for discovery, configuration, and telemetry collection.</p> </li> <li> <p data-renderer-start-pos="5624" data-local-id="04fd224fdc74">Demonstrates a high level of energy, results driven, and able to work under pressure with tight deadlines.<br>Technical Skills:</p> </li> <li> <p data-renderer-start-pos="5752" data-local-id="1480eb1b2e8a">High proficiency in Golang and at least one additional language (e.g., Python). You will write a lot of code in this role.</p> </li> <li> <p data-renderer-start-pos="5878" data-local-id="e5c76aaf8358">Strong Linux systems knowledge—you should live in Linux. Proficient with system administration, networking, storage, process management, log analysis, and troubleshooting at the OS level.</p> </li> <li> <p data-renderer-start-pos="6069" data-local-id="f5ff59240749">Experience with Grafana for building dashboards, alerting, and visualization of infrastructure metrics. Experience with Prometheus, InfluxDB, or similar observability platforms and building custom integrations/exporters.</p> </li> <li> <p data-renderer-start-pos="6293" data-local-id="c5d460a2f012">Experience with configuration management and infrastructure-as-code tools (SaltStack, Ansible, Terraform, or similar).</p> </li> <li> <p data-renderer-start-pos="6415" data-local-id="af2a131cf406">Solid understanding of networking concepts: L2/L3 protocols, VLANs, BGP, SNMP, and switch/router configuration (Arista, Cisco).</p> </li> <li> <p data-renderer-start-pos="6546" data-local-id="bc18cd0f5e5a">Experience with APIs and data integration—consuming vendor APIs, normalizing heterogeneous data sources, building data pipelines for metrics and reporting.</p> </li> <li> <p data-renderer-start-pos="6705" data-local-id="bc1fdd38052d">Experience with ClickHouse and MySQL—writing queries, designing schemas, and building tooling that reads from and writes to these databases.</p> </li> <li> <p data-renderer-start-pos="6849" data-local-id="ee4cd0f0d558">Experience with GitHub for version control, code review, CI/CD workflows, and collaborative development.</p> </li> <li> <p data-renderer-start-pos="6957" data-local-id="cec8c8c35b30">Demonstrated heavy use of AI tools (e.g., LLM-based coding assistants, AI-driven analytics) in a professional setting. You should already be using AI daily and be eager to push its application further.</p> </li> <li> <p data-renderer-start-pos="7162" data-local-id="fdc73214b70d">A compulsion to perform root cause analysis.</p> </li> <li> <p data-renderer-start-pos="7210" data-local-id="9ae310084452">Excellent written and verbal communication skills with the ability to work across a global engineering team.</p> </li> <li> <p data-renderer-start-pos="7322" data-local-id="19c48572c941">Extremely high personal standards for work quality.</p> </li> <li> <p data-renderer-start-pos="7377" data-local-id="42a99d216197">Reliable and predictable availability, including ability to work evenings and weekends as required.</p> </li> <li> <p data-renderer-start-pos="7480" data-local-id="ac3ec108535b">Bachelor's degree preferred.</p> </li> </ul><div class="content-pay-transparency"><div class="pay-input"><div class="description"><p><strong><span style="font-size: 14px;">Benefits</span></strong></p> <ul> <li style="font-size: 14px;"><span style="font-size: 14px;">Discretionary bonus eligibility </span></li> <li style="font-size: 14px;"><span style="font-size: 14px;">Medical, dental, and vision insurance</span></li> <li style="font-size: 14px;"><span style="font-size: 14px;">HSA, FSA, and Dependent Care options</span></li> <li style="font-size: 14px;"><span style="font-size: 14px;">Employer Paid Group Term Life and AD&amp;D Insurance</span></li> <li style="font-size: 14px;"><span style="font-size: 14px;">Voluntary Life &amp; AD&amp;D insurance</span></li> <li style="font-size: 14px;"><span style="font-size: 14px;">Paid vacation plus paid holidays</span></li> <li style="font-size: 14px;"><span style="font-size: 14px;">Retirement plan with employer match</span></li> <li style="font-size: 14px;"><span style="font-size: 14px;">Paid parental leave</span></li> <li style="font-size: 14px;"><span style="font-size: 14px;"><span style="font-size: 12px;"><span style="font-size: 14px;">Wellness Programs</span><br><br></span></span></li> </ul></div><div class="title">Annual Base Salary Range </div><div class="pay-range"><span>$150,000</span><span class="divider">&mdash;</span><span>$200,000 USD</span></div></div></div>

Apply at Jump Trading

Related jobs

Technical Project Manager

Jump Trading · Chicago, IL

Posted Sep 15 · Verified Sep 20

Network Engineer - Wireless

Jump Trading · Chicago, IL

Posted Sep 14 · Verified Sep 20

Data Center Technician

Jump Trading · Chicago, IL

Posted Sep 11 · Verified Sep 20

Senior Accounting Analyst | Corporate Accounting

Jump Trading · Chicago, IL

Posted Sep 9 · Verified Sep 20

HPC Data Center Infrastructure Planning Lead

Jump Trading · Chicago, IL

Posted Sep 2 · Verified Sep 20

Research Scientist/Research Engineer, Reinforcement Learning

Jump Trading · Chicago, IL

Posted Aug 26 · Verified Sep 20