Senior System Software Engineer
At Chicago UpDown, we are building breakthrough graph computing acceleration - software and hardware. With 1000x performance of today's cloud-based approaches, UpDown enables instant real-time mapping and analytics of trillion-node graphs. The core technologies were created as part of IARPA's AGILE program, and spun out from the University of Chicago and Purdue University.
We are looking for a talented System Software engineer, experienced with Linux and HPC, and capable with advanced AI tools to lead the design and enhancement of system software for a sophisticated, scalable global memory management and scale-out parallel cluster system (>10,000 nodes).
We are seeking a talented System Software Engineer to design, build, and optimize the memory and system software stack for our scalable, accelerator-based global memory system, UpDown.
Challenges include accelerator global memory (petabytes of unified virtual memory) and resource management in large-scale parallel systems (cloud and HPC-like). The technical challenges exceed the largest cloud and HPC systems today.
Key Responsibilities
Lead the design and implementation of a petabyte-scale unified virtual memory system to leverage novel hardware translation features and capabilities.
Global Memory and system library: System and application allocators & runtimes, including custom memory allocators and system-level runtime infrastructure to maximize throughput.
Task/Job Scheduling & Orchestration: Integrate and optimize HPC workload managers and job scheduling systems (such as Slurm, PBS, or custom schedulers) to efficiently allocate resources and orchestrate massively parallel jobs.
Kernel and Driver Engineering: Linux kernel and driver work to support and optimize the memory management and device driver subsystem, develop custom modules, and eliminate kernel-space bottlenecks.
Minimum Qualifications
Education: Master’s or Ph.D. in CS, CE, or related field.
Systems Programming Experience: 3+ years of hands-on experience writing production-grade, low-level code in C and C++ and/or Linux kernel programming.
Virtual and Memory Management Knowledge: Deep understanding of OS-level memory management, virtual memory, paging, and custom memory allocators.
Parallel Computing or HPC Experience: 3+ years in HPC/Cloud scalable systems software, writing/modifying software for HPC, supercomputer, and accelerator environments.
Preferred Qualifications
5+ years hands-on experience developing or modifying OS kernels (Linux), hypervisors, or low-level runtime libraries.
Strong understanding of the challenges associated with memory management in multi-tenant, globally distributed cloud infrastructure.
Advanced Interconnects: Demonstrated experience working with high-performance network and memory fabrics, such as RDMA (RoCE, InfiniBand), CXL, and NVLink.
Experience with graph workloads and algorithms (e.g., Graph500 BFS) on modern hardware.
To apply, send a cover letter and resume to recruiting@chupdown.com