Data center operations · Network infrastructure

Saeed Khan

Data Center Specialist

Keeping high-availability AI/HPC and cloud infrastructure up — from the fiber in the rack to the VLAN on the switch. 3+ years at NVIDIA and Amazon Web Services, troubleshooting Layer 1–3 connectivity, deploying and maintaining servers, switches, NVSwitches, storage and optics, and leading break-fix, capacity, and incident-response work.

  • 3+ yrs high-availability ops
  • Layer 1–3 troubleshooting
  • AI/HPC · Cloud
  • Fiber · Optics · DWDM
  • San Francisco Bay Area
LinkedIn
saeed@dc-ops — status
$
ifacerolestate · since
nvidia0Laboratory TechnicianData Center Operations · NVIDIAUP · Jun 2026
aws0DC Ops Technician L3Amazon Web Services · Santa Clara2023 → 2026
wgu0B.S. Computer ScienceWestern Governors UniversitySYNC · 2025
uptime 3y 8m layers L1–L3 optics OK tickets closed ✓
$
About

Hands-on, from Layer 1 up.

I'm a data center operations and network infrastructure technician with 3+ years supporting high-availability AI/HPC and cloud environments at NVIDIA and Amazon Web Services.

My work is hands-on: troubleshooting Layer 1–3 connectivity, deploying and maintaining servers, Ethernet switches, NVSwitches, storage, and fiber infrastructure, and leading break-fix, capacity, and incident-response work when availability is on the line.

I'm known for disciplined change execution, accurate documentation, and cross-functional coordination with engineers and vendors — the parts of the job that keep the next maintenance window boring.

San Francisco Bay Area

Disciplined change execution

Approved MOPs and SOPs followed to the letter, with post-work validation before anything touches production.

Accurate documentation

Cabling records, port maps, asset tags and ticket notes the next technician can actually use.

Cross-functional coordination

Engineers, vendors, stand-ups and incident bridges — status communicated early, blockers escalated fast.

Technical skills

The toolkit.

Four areas, all learned on the floor of live AI/HPC and cloud data centers.

Data Center & AI/HPC

Deploy · Maintain · Decommission
  • Rack & stack
  • Hardware bring-up
  • Deployment
  • Decommissioning
  • Break-fix
  • Preventive maintenance
  • Capacity planning
  • Asset lifecycle

Networking

Layer 1–3 · Copper · Fiber
  • Layer 1–3 troubleshooting
  • Copper / fiber
  • Optics
  • Cable tracing
  • Switch-port mapping
  • VLAN / IP validation
  • Cisco
  • Juniper MX / QFX / EX / PTX
  • Ciena WaveServer
  • DWDM

Hardware & Systems

Servers · GPUs · Switches · Storage
  • Linux
  • Servers
  • GPUs
  • Ethernet switches
  • NVSwitches
  • Storage
  • NICs
  • Power supplies
  • Console / OOB access
  • Diagnostics
  • Component replacement

Operations & Controls

Tickets · Change · Vendors
  • Ticketing
  • Change management
  • MOP / SOP execution
  • Incident escalation
  • RMAs
  • Vendor coordination
  • Inventory & asset records
  • Documentation
Cover letter

Why I'd be a strong addition to your team.

I'm a data center operations and network infrastructure technician with more than three years supporting high-availability AI/HPC and cloud environments at NVIDIA and Amazon Web Services. I'm writing to express my interest in joining your team, where I can bring hands-on experience keeping critical infrastructure up, well documented, and ready for what comes next.

At NVIDIA, I troubleshoot and resolve hardware, network, and physical infrastructure issues across enterprise AI and HPC environments. Day to day, that means isolating Layer 1 faults by validating copper and fiber paths, verifying switch-port mappings and link state, and replacing faulty optics or components; installing, cabling, and configuring servers, Ethernet switches, NVSwitches, and storage to engineering standards; and running console and out-of-band diagnostics during deployments and break-fix work. Every job ends with post-installation quality checks and accurate asset records before equipment is handed off to production.

Before that, I spent three and a half years at AWS in Santa Clara, where I was promoted to Data Center Operations Technician L3. I led complex capacity work, rackdowns, and high-severity incidents across multiple points of presence, troubleshooting Cisco, Juniper MX/QFX/EX/PTX, Ciena WaveServer, and DWDM infrastructure. I diagnosed connectivity across OSI Layers 1–3, replaced optics, installed and traced fiber, managed RMAs, and owned tickets through verification and closure. I also mentored technicians and built a training resource for interns, because the fastest way to raise availability is to raise the whole team.

What I'm known for is disciplined change execution: approved MOPs and SOPs followed to the letter, clear status during stand-ups and incident bridges, blockers escalated early, and documentation the next technician can actually use. I'm currently completing a B.S. in Computer Science at Western Governors University, building on an A.A.S. in Information Technology with an emphasis in network administration.

I would welcome the chance to discuss how my experience can support your data center and network operations. Thank you for your time and consideration.

Professional experience

Where I've kept the lights on.

Laboratory Technician, Data Center Operations

NVIDIA · Enterprise AI & HPC environments
Jun 2026 — Present
  • Troubleshoot and resolve hardware, network, and physical infrastructure issues across enterprise AI and high-performance computing (HPC) environments, minimizing downtime and restoring service availability.
  • Diagnose Layer 1 connectivity faults by validating copper and fiber cabling, tracing cable paths, verifying switch-port mappings and link status, and isolating faulty optics, ports, or network components.
  • Install, rack, stack, cable, configure, and replace servers, Ethernet switches, NVSwitches, storage systems, and other data center hardware in accordance with operational and engineering standards.
  • Perform console and out-of-band troubleshooting, controlled power cycles, hardware diagnostics, component replacement, and post-maintenance validation during deployments and break-fix operations.
  • Complete post-installation quality checks on rack placement, labeling, cable routing, port mappings, link state, and asset records before equipment handoff or production use.
  • Receive, stage, and inspect hardware; verify serial numbers, asset tags, MAC addresses, rack/U locations, and inventory updates; coordinate parts replacement, returns, and RMAs.
  • Execute approved maintenance and change procedures, track deliverables and risks, communicate status during DC Ops stand-ups or incident bridges, and escalate blockers to engineering teams.
  • Monitor and resolve tickets with detailed troubleshooting, root-cause, validation, and resolution notes; maintain cabling records, port documentation, work instructions, and operational procedures.
  • Coordinate capacity expansion, preventive walk-throughs, decommissioning, vendor escorts, and remote-hands support; maintain organized racks and shared spaces, verify cable management and airflow, and report infrastructure risks before production impact.

Data Center Operations Technician

Amazon Web Services · Santa Clara, CA
Jan 2023 — Jun 2026
Promoted to Data Center Operations Technician L3 · Jul 2024
  • Led complex capacity work, rackdowns, high-impact network incidents, and high-severity production events across AWS data centers.
  • Supported multiple points of presence (POPs) and troubleshot Cisco, Juniper MX/QFX/EX/PTX, Ciena WaveServer, and DWDM infrastructure.
  • Led infrastructure and capacity projects at additional POP sites, coordinating site readiness, vendor work, fiber/cabling, hardware deployment, and engineering closeout.
  • Diagnosed connectivity across OSI Layers 1–3, including fiber and optic faults, interfaces, VLANs, switching, and IP reachability; assisted teams with coordinated Layer 4–7 troubleshooting.
  • Assisted engineering teams with rack bring-up and power-up activities, supporting hardware installation, cabling, power validation, network connectivity checks, and post-deployment verification.
  • Configured switches, replaced optics, installed and traced fiber, validated interface status, managed RMAs, and restored network service during maintenance and incident response.
  • Performed server break-fix and component replacement involving CPUs, GPUs, RAM, NICs, power supplies, fans, HDDs, SSDs, M.2 drives, cabling, and system boards.
  • Owned operational tickets through verification and closure, coordinated engineers and vendors during change windows and incident bridges, documented work to service standards, and performed post-work quality checks to identify risks before customer impact.
  • Mentored technicians, created an intern training resource, partnered with engineering teams, and contributed to operational metrics, service goals, and high-availability support.
Education

Still learning, on purpose.

Bachelor of Science, Computer Science

Western Governors University
May 2025 — Present

In progress. Relevant coursework: Data Structures and Algorithms, Discrete Mathematics, ITIL Methodologies, Networking Fundamentals.

Associate of Applied Science, Information Technology

Heald College · Stockton
2012 — 2014

Emphasis in Network Administration.

Contact

Let's connect.

Open to data center and network infrastructure roles in the San Francisco Bay Area. The fastest way to reach me is email.

linkedin.com/in/saeedkhan1994
Available for new opportunities San Francisco Bay Area
Copied