Week 18 · lesson

Maintenance Protects Availability

A mouse fails at Station 04 during finals. Annoying — but did the problem begin during finals, or had that mouse been double-clicking, disconnecting, acting strange for two weeks with nobody reporting it? If it's the second one, "surprise" isn't really the right word.

Failure has warning signs

Some failures happen suddenly. Others develop over time: an intermittent connection, an unusual noise, low storage, a damaged cable, a repeated software error, battery degradation, a peripheral input problem. A warning sign is evidence that a system's condition may be changing — not proof of a specific root cause. Same evidence-first habit from Week 16, just applied earlier in the timeline, before the failure happens instead of after.

Availability, made concrete

If a lab has eight player stations and all eight work, availability for the event is good. One completely fails — now there are seven usable stations. If the event requires exactly eight, that one failure suddenly affects the whole scheduled match, not just one seat. Context decides how much a single failure actually matters.

Spares change the situation

If one mouse fails but the lab has three tested spares, the failure still matters, but it probably doesn't stop the event — the spare reduces the operational impact almost to nothing. This is exactly why the Week 15 inventory work mattered: knowing what backup exists changes how urgent a failure actually is.

Maintenance protects function, not objects

Nobody maintains a mouse because the mouse deserves it — it's maintained because it supports player input, which supports gameplay, which supports match availability. Trace it: a frayed Ethernet cable creates a possible connection instability, which creates a player-station reliability risk, so it gets replaced before the event, then verified. That's preventive maintenance, and the function it protects is the whole reason it's worth doing.

Preventive maintenance

The equipment still works, but you act because evidence suggests maintenance now may prevent a later problem: replacing damaged cable, cleaning approved equipment, clearing approved storage, completing a required software update before event day, confirming spare equipment condition ahead of time.

Corrective maintenance

Failure has already happened. A dead mouse gets replaced. A damaged game installation gets repaired or reinstalled through an approved process. Corrective maintenance restores the expected function after something has already broken.

Condition-based maintenance

The interesting middle ground. A headset works most of the time, but students report the microphone disconnecting occasionally. Nothing has completely failed — but ignoring it probably isn't right either. The observed condition justifies investigation or scheduled attention, even though there's no full failure to point to yet.

Maintenance isn't random replacement

"This keyboard is three years old" doesn't by itself prove it needs replacing — it might work perfectly, while a six-month-old keyboard sitting next to it is already failing. Use condition, function, evidence, and requirement to decide, not a device's age alone.

Maintenance records

Station 06 has a microphone problem; a technician swaps the headset; everything works. Is that the end of it? Almost — record the date, asset, symptom, action, result, and follow-up. For example: Headset-06, microphone input dropped twice during practice, replaced with a tested spare, input verified in the voice application, original headset marked for further inspection. That record is useful the next time anyone looks at that station.

Without records, the same issue happening three times just feels random. With records, a pattern can become visible — same station, same application, same port, same type of failure — which is information a single incident on its own can't give you.

Investigation: maintenance or not?

For each case, choose monitor, preventive, condition-based, corrective, or escalate, and explain why.

A mouse that completely fails: corrective, straightforward. A damaged Ethernet cable jacket where the connection still works: preventive attention makes sense before it becomes a real failure. A computer with normal storage and no reported problems: probably just monitor — not everything needs intervention. A headset microphone that disconnects once every few practices: condition-based investigation. A network switch serving half the lab that repeatedly loses connectivity: students can inspect and document, but freshmen shouldn't be randomly reconfiguring production network hardware — escalate.

Maintenance boundaries

Students can safely perform actions like visual inspection, approved peripheral swaps, documentation, verifying software state, checking available storage, reporting, and testing after an approved change. Other work — internal hardware repair, managed network reconfiguration, vendor-level fixes — belongs to an instructor, IT staff, an authorized technician, or a vendor. Good maintenance planning includes knowing when not to touch something yourself.

The maintenance chain

Observe, record, assess, act or escalate, verify, follow up. Verify shows up again — Week 17 refuses to disappear, because verification matters at every layer of a system, not just the one you happened to learn it in.

Maintenance investigation: five reports

Report A. Station 02 mouse double-clicks occasionally.

Report B. Broadcast workstation has 4 GB free storage remaining.

Report C. Station 07 monitor works normally.

Report D. One network cable has a damaged connector.

Report E. Switch-B has disconnected several stations twice this month.

For each: current condition, function affected, maintenance type, evidence, action or escalation, and verification.

Don't fix what you don't understand

"Low storage" doesn't mean "delete everything." Ask what's actually using the storage, what's approved for removal, whether recordings are supposed to be archived somewhere first, and who owns the files. Maintenance should never create a new failure while solving the original one.

Before you leave

Complete:

  1. A system can need maintenance even while it still works because __________.
  2. One piece of evidence that should trigger attention before total failure is __________.

Be specific — vague answers don't teach anyone reading them anything new.

Source note

This lesson applies Middle Township Unit 5 concepts involving maintenance, repair, dependability, availability, system evaluation, and technology constraints.

The maintenance-type model, availability analysis, maintenance-record structure, student safety boundary, and fictional maintenance cases are original Robotnix Academy instructional components.