Data center technician interviews check whether you can be trusted alone on a floor full of other people's equipment. Expect a few questions on why you want shift work, hands-on questions about racking servers, copper and fibre cabling, swapping failed parts, power feeds and cooling, and plenty of what-would-you-do scenarios where the right answer is to stop and check before touching anything. Each question shows what the interviewer is really listening for, a shape for your answer, and a short answer you could say out loud. Swap in your own stories and your own site's rules before the day.
Search all questions by round, difficulty and level, or save the ones you want to practise.
Start: the first time you took hardware apart or built something, kept short.
Proof: a job, course or project where you worked with real equipment.
Why here: what you like about the floor, such as fixing physical things that keep services running.
"I started building my own PCs as a teenager, and later I did a year on a help desk where I was always the one who volunteered to swap parts or run cables. That's where I noticed I prefer work I can see and touch. I did a hardware and networking course, and for my final project I set up a small rack with a switch, two old servers and a patch panel, labelled everything and documented it. A data center is that same work at a bigger and more careful scale. I like that it's methodical, that there's a right way to do each job, and that when I get it right, services people depend on just keep running."
Saying you want the job as a quick way into software or cloud work, with no interest in the hardware itself.
Honest yes: say plainly that you can work the pattern, and how you manage sleep and life around it.
Evidence: any past shift work or on-call.
Your questions: two or three practical things, such as how escalation works at night and how many people are on site.
"I'm genuinely fine with it. In my last job I worked a rotating pattern in a warehouse, so I know I need a fixed sleep routine on night weeks and I plan around it. What I'd want to know about your site is mostly practical. How many technicians are on a night shift, and who do I call when something is beyond me, especially for facilities problems like power or cooling? What does the handover between shifts look like? And what are the most common tickets, so I know what I'll be doing most. Knowing those things early helps me be useful on my own sooner."
Agreeing to any shift pattern without thought, or asking only about pay and time off.
Done: the kit you've physically installed, swapped or cabled.
Read about: the areas you know in theory only.
Plan: how you'd close the most important gap in the first months.
"I've racked and cabled rack servers, replaced disks, memory and power supplies, and made up and tested copper patch cables. I've plugged in fibre patch leads and optics, but I've never terminated or spliced fibre myself, and I haven't used a light meter on the job, only in a training lab. On the power side, I understand how UPS, generators and PDUs fit together, but I've never been part of a live generator test or maintenance event. So my biggest gap is facilities work. I'd want to shadow the facilities team during their planned tests and learn exactly where my job stops and theirs starts."
Claiming experience with everything, then struggling when asked for one concrete detail.
Before: check the ticket, the rack elevation, the U position, power and ports.
Install: inspect, fit rails, lift with two people, mount, cable neatly.
Power and prove: A and B feeds, check lights and health.
Close out: labels, asset records, photos and ticket notes.
"First I check the ticket against the rack elevation, so I know the exact rack, the U position, which PDU outlets and which switch ports it gets. At the dock I check the box for damage and match the serial to the order. I fit the rails at the right height, and for anything heavy I get a second person to lift, never balance it alone. Once it's in, I cable it along the cable management so nothing blocks the fans or the next server, and I plug the two power supplies into the A and B PDUs. I power it on, check the lights and the management controller for faults, then label both ends of every cable, update the asset records, and close the ticket with photos."
Skipping the ticket and records check, or lifting a heavy server alone to save time.
Core: single-mode has a very thin core for long runs; multimode has a wider core for short runs.
Light source: single-mode optics use lasers at longer wavelengths; common multimode optics use cheaper short-wavelength lasers.
On the floor: jacket colour as a hint, the label and the records as the proof, and optics that match.
"Single-mode fibre has a very thin core, around nine microns, so light travels in basically one path. That's why it can go kilometres, and it's used with long-reach optics. Multimode has a much wider core, usually fifty microns in modern cable, so light takes many paths and it's limited to shorter runs, which is fine inside a data hall. On the floor, the jacket colour is a good hint: single-mode is usually yellow, newer multimode is usually aqua, and older multimode is often orange. But colour is a convention, not a guarantee, so I also read the cable print and check the records. The key thing is the optic must match the fibre type, or the link won't come up or will be unreliable."
Saying you can tell the type from the connector alone, or that any optic works on any fibre.
Inspect first: use a fibre scope on the end face.
Clean: a proper click cleaner or lint-free wipe with the right fluid, never a shirt or breath.
Inspect again: only plug in when it's clean, and cap anything unused.
Safety: never look into a fibre or optic directly.
"The rule I follow is inspect, clean, inspect. I look at the end face with a fibre scope first, because even a new patch lead out of the bag can be dirty. If there's dust or a smear, I clean it with a click cleaner or a lint-free wipe and the proper cleaning fluid, then check it again under the scope. I only plug it in when it's clean, and I do the same for the port if it looks suspect. The fuss is because a tiny bit of dust on the core causes loss and reflections, and if you mate two dirty connectors you can scratch both end faces permanently. I also never look into a fibre with my eye, since the light can be invisible and still harmful."
Blowing on the connector or wiping it on clothing, or never having heard of a fibre scope.
Scope: how much equipment, how many people, how long.
Method: elevations, cable plan, labels made in advance, a checklist.
Checks: how you verified each rack before calling it done.
Result: what went well and one thing you'd change.
"At my last company we brought in four racks of new servers, around eighty machines, over two weeks with three of us. Before anything arrived, we agreed the rack elevations and a cable plan, and printed labels for every cable and port in advance so nobody was inventing names on the floor. We worked in a fixed order: rails, servers, power, then data, one rack at a time. Each rack had a checklist, and a second person checked it against the plan, including power supply status and link lights on every port. We found about half a dozen crossed cables that way before handover. What I'd change is scanning serials into the asset system at the dock instead of at the end, which cost us a day."
Describing speed only, with no mention of how the work was checked or recorded.
Found: what didn't match and how you noticed.
Checked: how you worked out which side was right.
Fixed: records, labels, and who you told.
Prevented: a habit or change that stops it recurring.
"I was sent to replace a disk in a server that the records put in U20 of a rack. U20 held a different model entirely. I found the right server by serial three units lower, and there were two more devices in that rack that weren't in the records at all. I did the disk job once I'd confirmed it with the owner, then walked the whole rack, noted every device with serial and position, and sent it to the team that owns the asset system with photos. It turned out a move had been done months earlier without the records being updated. After that we added a rule that no move ticket closes until the records are updated and checked by a second person."
Doing the job and leaving the wrong records as they were because it wasn't your ticket.
Confirm: right server by asset tag and serial, right slot by fault light and management controller.
Swap: hot-swap only if the server supports it, matching replacement part.
Verify: the rebuild starts and the customer or owner is told.
Old disk: follow the data destruction or return policy.
"I start with the ticket and confirm the server by its asset tag and serial, not just the rack position. Then I confirm the drive: the slot should show a fault light, and I check the management controller or ask the owner to confirm the slot and the disk's serial. That matters because the array is already degraded. In something like RAID 5, pulling a healthy disk by mistake takes the whole array down. If the server supports hot-swap, I pull the failed disk, wait a few seconds, and insert the matching replacement. Then I check that the rebuild has started and tell the requester. The old disk goes through the site's data destruction or return process, never into a bin or a drawer."
Pulling a disk based only on the slot number in the ticket, or treating the old disk as ordinary waste.
Your change first: reseat the new modules and check the slots used.
Rules: memory population rules and matching part numbers.
Evidence: front panel codes, lights and the management controller's event log.
Isolate: minimum config, swap known-good parts, then escalate with findings.
"The first suspect is whatever I just touched. I power it down, pull the power cords, put on my wrist strap, and reseat the new modules properly, checking the latches closed on both sides. Then I check the slots against the server's memory population guide, because most servers want modules in a set order per processor, and mixed part numbers or speeds can stop it booting. Next I read the evidence: the front panel or diagnostic lights, any error codes, and the management controller's event log, which often names the exact slot. If it's still stuck, I go to a minimum config, one module per processor in the first slots, and add the rest back a few at a time to find the bad module or slot. I note everything so the vendor call is quick."
Swapping random parts without reading the event log, or forgetting ESD protection while handling memory.
What it is: a small separate computer on the motherboard with its own network port.
What it gives: power control, remote console, sensor readings and a hardware event log.
Why it matters: it works when the operating system is dead.
"It's a small computer built into the server's motherboard, often called a BMC, and vendors give it their own names. It usually has its own network port on a separate management network, and it runs as long as the server has power, even if the operating system has crashed or the server is switched off. For me it's the first place to look. It shows fan, temperature and power supply health, and it keeps a hardware event log that usually names the failed part or slot. It also lets someone power cycle the server or use a remote console without standing in front of it, which is why its port needs to be cabled and labelled correctly on day one."
Confusing it with the operating system's remote desktop, or not knowing a hardware event log exists.
Stop: pull nothing yet.
Gather facts: slot numbering, management controller view, disk serials.
Confirm: with the requester, reading back what you see.
Record: the mismatch in the ticket, so the next job is safer.
"I stop and pull nothing. On a degraded array, pulling the wrong disk can take the whole thing down, so a guess is the worst option. First I check how this model numbers its slots, because some management tools count from zero while the labels on the front start at one, and that would turn the ticket's slot 3 into the bay marked 4. Then I look at the management controller to see which slot and serial it reports as failed. After that I contact the requester, tell them exactly what I see, read back the serial of the failed disk, and ask them to confirm from their side, maybe by lighting the locate LED on the right drive. Only once we agree on one drive do I swap it. I write the mismatch in the ticket either way."
Pulling the drive in slot 3 because the ticket says so, or pulling slot 4 without telling anyone.
Symptom: what people saw and why it was hard.
Method: logs, patterns, swapping one thing at a time.
Cause: what it turned out to be.
After: how you stopped it happening again.
"We had a server that rebooted itself every few days with nothing useful in the operating system logs, and the owner had already had its memory replaced once. I pulled the management controller's event log and saw two things. One power supply had been reporting failed for weeks, but that alert went nowhere. And just before each reboot, the other supply logged a loss of input power. So the server was running on one leg, and that leg kept dropping. I checked the times against the access log, and every reboot lined up with someone working at the back of that rack. The cord for the good supply was slightly loose in the PDU, and opening the rear door nudged it. We replaced the failed supply, fitted locking cords and got the supply alerts sent to the right team. Now I always read the hardware log before anyone swaps parts."
A story where parts were swapped at random until the problem went away.
Path: utility, switchgear and transfer switch, UPS, distribution, rack PDUs, server power supplies.
Generator: takes over when utility power fails.
UPS: bridges the gap and cleans up the power.
Redundancy: separate A and B paths down to the rack.
"Power comes in from the utility to the site's switchgear. There's a transfer switch that can move the load onto the generators if the utility fails. The generators take a short while to start and pick up the load, so the UPS sits in between: it runs off batteries or a flywheel for those seconds, and it also smooths out spikes and dips. From the UPS, power goes through floor distribution units and breaker panels out to the rack PDUs, and from there into the servers' power supplies. In a well-built site there are two separate paths, an A side and a B side, so a failure or maintenance on one path doesn't take down servers that are plugged into both."
Thinking the UPS runs the site for hours, or not knowing the A and B paths are separate.
Plugging: each power supply to a different feed, never both to the same PDU.
Load: each feed must carry the full rack load if the other fails, within the site's limit.
Checks: both supplies healthy, cords labelled, load readings reviewed.
Records: what's plugged where, kept current.
"Two power supplies only help if they're on truly separate feeds, so first I check each server has one cord in the A PDU and one in the B PDU, labelled, and that both supplies show healthy in the management controller. A supply that's failed or unplugged means that server is running on one leg without anyone knowing. The bigger trap is load. When the A side fails, everything shifts to B, so the combined load of both feeds has to fit on one circuit within the limit the site allows. If each feed is sitting just under its limit, the rack looks fine every day and trips the moment one side drops. So I watch the PDU readings and flag any rack where the total is creeping past what one feed can hold."
Saying each feed can be loaded to its own limit, forgetting that one feed must carry everything during a failure.
Layout: equipment fronts face the cold aisle, exhausts face the hot aisle.
Why: keep cold supply air and hot exhaust from mixing.
Mistakes: missing blanking panels, gear mounted backwards, open containment doors, badly placed floor tiles, unsealed cable holes.
"Racks are set up in rows so the fronts of the equipment face each other across a cold aisle, where the cooling units push cold air, and the backs face each other across a hot aisle, where the exhaust goes back to the cooling units. The whole point is to stop hot and cold air mixing. The mistakes are usually small. Leaving empty U spaces without blanking panels lets hot air loop back to the front. Mounting a switch backwards, or with the wrong fan direction, blows hot air into the cold aisle. Propping a containment door open, moving a perforated tile to the wrong place, or leaving a cable hole unsealed all leak air. I fit blanking panels whenever I remove something."
Not knowing which side of a server takes air in, or treating blanking panels as cosmetic.
Measure: current load on both feeds, not just one.
Estimate: what the new servers will draw, from the spec or similar servers.
Failure case: could one feed carry everything if the other drops?
Raise it: capacity team or customer, with options.
"I wouldn't install and hope. First I'd check the current readings on both the A and B feeds, and the site's allowed limit for those circuits. Then I'd estimate what the new servers will draw, using the vendor's figures or readings from the same model elsewhere, since nameplate ratings are usually well above real draw. The test that matters is the failure case: if one feed drops, can the other carry the whole rack including the new servers without tripping? If it can't, I raise it with the capacity or planning team and the customer before doing anything, and offer options like another rack with headroom or an extra circuit. I'd note the numbers in the ticket so the decision is on record."
Installing anyway because there are free outlets, or only looking at one feed's reading.
Right place: correct server port and switch port as per the ticket and patch records.
Cable and optic: seated, undamaged, right type, fibre clean and polarity right.
Prove it: link lights, a known-good cable, a tester or light meter.
Hand over: what you checked, what you saw, what's left.
"First I check it's patched where the ticket says, on both ends, including any patch panels in between, because a wrong port is the most common cause. Then I look at the link lights on both sides. For copper, I reseat the cable and try a known-good one, or test it. For fibre, I check the optic type matches the fibre and the switch side, inspect and clean the ends, and check polarity, since transmit and receive swapped means no link. If I have a light meter, I check that light is actually arriving. If all of that is right, I hand it to the network team with the exact ports, what I tested and the results, so they can check whether the switch port is enabled and set up correctly."
Logging into the switch to change settings outside your role, or escalating without having checked the cable.
Network: list interfaces and their state, then check one link's speed.
Disks: list the disks the system sees, knowing a RAID card hides the physical ones.
Errors: recent kernel messages about hardware.
Identity: the serial number to match the ticket.
"I keep it to read-only checks unless the owner asks for more. For network, ip -br link shows each interface and whether it's up, and ethtool on one interface tells me if a link is detected and at what speed. For disks, lsblk shows what the system can see. Behind a hardware RAID card it only shows the array as one disk, so for a rebuild I check the controller or the management controller instead. For errors, recent kernel messages from dmesg often show disk or memory faults with a timestamp. And I read the system serial number to be sure I'm on the server the ticket names. Anything beyond that, like changing configuration, I leave to the system owner."
ip -br link # interfaces and up/down state
sudo ethtool eno1 # link detected, speed, duplex
lsblk # disks the OS can see
sudo dmesg -T | tail -n 50 # recent kernel messages
sudo dmidecode -s system-serial-number # match the ticket
Restarting services or editing network config on a customer's server just to make a check pass.
lsblk doesn't show it. What might explain that?What it is: a static discharge from your body into a component.
Why it matters: damage can be instant or hidden until later.
Habits: wrist strap, antistatic bags and mats, hold parts by the edges.
"Electrostatic discharge is the static charge that builds up on your body and jumps to a component when you touch it. You might not feel it, but it can damage memory, processors or boards, and sometimes the damage doesn't show until the part fails weeks later, which is worse. So I wear a wrist strap clipped to a proper ground point, usually the server chassis or a mat, whenever I open a server. I keep new and removed parts in their antistatic bags, put them down on an antistatic mat rather than the top of a server, and hold them by the edges, never by the contacts or chips. It takes seconds and it's cheaper than a mystery fault."
Saying you just touch the metal case once and carry on, or that ESD only matters in dry weather.
Approach: calmly ask who they are and who they're with.
Verify: badge, escort and access record, through security.
Contain: don't leave them alone in the hall; close the door.
Report: log it even if it turns out innocent.
"I'd approach them politely and ask who they are and who they're working with. Most of the time it's a contractor whose escort stepped away, but I can't assume that. I'd ask to see their badge and I'd call security straight away to check the access record, staying with them or keeping them in sight rather than walking off. If they can't be verified, security takes over and escorts them out. I'd also close and check the propped door, since that's a breach on its own, and look at whether anything nearby was touched. Then I'd log the whole thing, even if it turns out to be innocent, because the process failed somewhere and that's worth fixing."
Letting them carry on because they look like they belong, or confronting them aggressively on your own.
Pause: stop at a safe point and don't power off or remove anything more.
Check: is the ticket wrong, the records wrong, or am I at the wrong device?
Escalate: the change owner decides; you don't improvise.
Fix records: whatever the answer, update them.
"I stop at a safe point. If I've only removed labels or unused cables, I leave things as they are and touch nothing else. A decommission is hard to undo, and pulling the wrong live server is a serious outage. Then I check whether I'm at the right rack and U position, whether the device has been moved, and what the asset records say. I call the change owner, explain exactly what I see, and let them decide whether to continue, correct the ticket, or back out. If the window runs out, we back out and reschedule. Whatever the outcome, the records need fixing so the next person doesn't hit the same trap."
Carrying on because the rack position matches, or quietly fixing the ticket yourself afterwards.
Situation: what you were doing and the mistake, plainly.
Response: how fast you raised it and helped fix it.
Change: the specific habit you use now.
"In my last job I was tidying cables at the back of a rack after an install, and I unplugged a patch lead I thought was left over from the old server. It was actually the uplink for a small switch serving another team's test lab. They lost access for about ten minutes. As soon as I saw the link light on the switch go out, I plugged it back in, then called our operations desk and told them exactly what I'd done instead of waiting for a ticket. Nothing was lost, but it shook me. Now I never unplug anything that isn't named in the ticket, and if a cable is unlabelled I trace it and check the records before it comes out."
Choosing a story with no real mistake, or one where you kept quiet and hoped nobody noticed.
Situation: the task and the shortcut on offer.
Choice: what you did and why.
Outcome: what happened, and whether the procedure mattered.
"We had to replace a failed power supply on a storage array late on a Friday. The procedure said to confirm the healthy supply was carrying the load and get the owner's go-ahead before pulling the failed one. A colleague said it was hot-swap, just pull it and go home. I still checked, and the management page showed the other supply was reporting a warning too. If I'd pulled the failed one, the array could have lost power entirely. I stopped, called the owner, and we got the vendor involved to replace both properly the next morning. It took an extra half hour of my evening, but it showed me why the checks are written down."
A story where the procedure was followed only because someone was watching.
Meaning: anyone can stop work that looks unsafe, without blame.
Everyday: lifting, ladders, trip hazards, working near live power only if qualified.
Speaking up: calmly, in the moment, then through the proper channel.
"To me it means anyone on the floor can stop a job that looks unsafe, and nobody gets blamed for asking. Most risks are everyday ones: lifting heavy servers alone, standing on a chair instead of a proper step, cables across walkways, or someone opening electrical panels they're not qualified to work on. And yes, I'd stop a senior colleague. I'd do it calmly and in the moment, something like, hang on, can we get a second person for that lift? If they brushed it off and it was dangerous, I'd stop the work and raise it with the shift lead. Seniority doesn't change what a falling server or live power can do to someone."
Saying you'd let a senior person carry on because they know best, or treating safety as the facilities team's job only.
Check: go and look, and read the cold-aisle inlet temperatures, not just the hot aisle.
Cause: cooling units in that zone, open containment, blocked airflow.
Escalate: follow the runbook, call facilities on-call and tell the operations centre.
Record: times, readings and actions for the handover.
"First I'd go to the row and see it for myself, with the monitoring open on my tablet or phone. A hot aisle is meant to be hot, so the number that really matters is the inlet temperature on the cold side. If inlets are fine and it's one sensor, it may be a sensor fault, which I'd still log. If inlets are rising, I check the cooling units for that zone for a fault, a stopped fan or an alarm, and look for simple causes like a containment door propped open or panels removed during a job. I'd call facilities on-call as the runbook says and tell the operations centre early, rather than waiting until I'm sure. I wouldn't reset cooling units or shut down customer equipment myself unless the emergency procedure tells me to. And I'd note every reading with the time."
Silencing the alarm and waiting to see, or resetting facilities equipment or powering off customer servers on your own.
Mindset: rounds exist to catch small things before they become big ones.
Habits: a checklist, looking for change, noting readings.
Example: something small you caught on a routine walk.
"I remind myself that rounds are how small problems get caught before they become 3 am incidents. What keeps me sharp is looking for change rather than looking for disasters. I follow the checklist, but I also notice things like a new noise from a fan, an amber light that wasn't there yesterday, a floor tile out of place, or a door that doesn't close properly. I write down readings instead of just ticking a box, because it makes me actually read them. On one walk I noticed a faint smell near a PDU, reported it, and the electrician found a loose connection that was starting to heat up. That was a quiet week too."
Admitting you tick the checklist from memory, or saying nothing ever happens so rounds are pointless.
Who: confirm the requester is authorised for that customer.
What: match cage, rack, U position, asset tag and serial to the ticket.
How: agree graceful versus hard power cycle, and confirm on the phone before acting.
Report: time done, what you saw, photos if useful.
"First I check the request came through the proper channel and the person is on that customer's authorised contact list. Then I find the device and match everything: the cage, rack, U position, asset tag and serial number. Racks often hold near-identical servers, so a position alone isn't enough. Next I confirm the method. A graceful restart from the console is very different from pulling power, and if they want a hard power cycle I ask them to confirm it can't be done remotely first. Ideally I have them on the phone, I read back the serial, and only act when they say go. Afterwards I note the exact time, the lights I saw on boot, and update the ticket."
Power cycling the server that sits in the U position named on the ticket without checking its serial.
Impact first: a live outage usually beats planned work.
Commitments: a scheduled change has a window and people waiting.
Delegate or defer: the delivery can usually wait or go to someone else.
Tell people: update everyone whose work moves.
"The server that's down is hurting someone right now, so it comes first unless the runbook says otherwise. The fibre patch is booked in a change window, so I'd call the change owner straight away, explain I'm on an outage, and ask whether they can hold for a few minutes or whether another technician can take it. The delivery is the least urgent; I'd ask security or a colleague to sign it in and keep it secure until I'm free, as long as nothing is left unattended on the dock. Then I work the outage and give updates as I go. If I'm the only person on site and it keeps happening, I'd raise the staffing gap with my lead afterwards."
Doing jobs in the order they arrived without telling anyone, or leaving a delivery open on the dock.
Situation: the problem and who was on the other end.
How you described it: exact labels, lights, positions, photos.
Result: the decision they made and how it went.
"A customer called because a switch in their rack had gone unreachable. When I got there, the switch was on, but one power supply light was off and the other was amber. Saying it looks broken wasn't going to help them, so I told them the rack, U position and serial, described each light by its label and colour, and sent photos of the front and back, which our site allowed for a customer's own rack. I also noticed its uplink cable was hanging out of the patch panel. We reseated the uplink together on the phone, it came back, and they opened a vendor case for the power supply using my photos. They told me later the photos saved them a site visit."
Vague descriptions like 'some lights are red', or acting on your own guess without the customer agreeing.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your resume and notes are never stored on our servers. It stays out of screen share on every plan; only you can see it.