DELTA-SIERRA · SPACE ACADEMYBack to the Mars Library
MODULE 15 · Progressive training: understand, calculate, verify.

Mission team and operations: decide far from Earth

Training Mars operations organisation linking crew, automation, Earth, procedures and mission log.
Mars operations distribute authority among automation, local crew and Earth expertise according to the time available.
Mars field team coordinating sampling, rover, drone and local instrumentation.
Conceptual visualisation of a multi-actor science operation. The team must coordinate humans, robotics, communications, safety, available time and sampling priorities, with procedures that can safely stop the activity if a critical resource degrades.
Pressurised vehicle connected to a Mars habitat through a closed transfer interface.
Conceptual visualisation of a pressurised logistics docking. Crew or cargo transfer is safe only after alignment, locking, leak checks and pressure equalisation; the image should not be read as an open doorway directly to the Martian atmosphere.
Crewed Mars convoy travelling along a rough route far from the base.
Conceptual visualisation of mobile operations: every traverse combines route, energy, communications, local weather, towing capability, degraded return and abort decisions. The operation is a system, not merely a drive.

Many Earth-orbit operations can rely heavily on a real-time ground team. Mars breaks that assumption: light-time alone introduces minutes of delay and outages are possible. Mission teams need intent-based work, complete information packages, adaptable procedures and explicit local authority. This module treats operations as an engineered subsystem.

1. A role needs both authority and information

Field laboratory analysing Martian cores and samples.
Science operations must preserve sampling, identity, contamination control and traceability while sharing crew time with life-critical tasks.

“The commander decides” is incomplete. Decides what, within what time, using which data and up to what risk? Responsibilities can be distributed across immediate safety, vehicle systems, medical work, navigation, science and maintenance. A small crew may combine roles, but ownership should remain explicit.

An operational role is more than a title. It combines decision authority, access to information and responsibility for handover. If two people both believe they own an urgent action, they may act in parallel; if each assumes the other owns it, nobody acts. Procedures therefore state who may stabilise the system, who authorises irreversible action and who maintains the timeline. Authority can change with mission phase and communications delay. During a Mars emergency, the local crew necessarily has more autonomy than a crew near Earth. That autonomy must be prepared through boundaries, thresholds and training rather than improvised after the anomaly begins.

2. Latency turns conversation into delayed messaging

Team performing sampling, measurements and geological documentation on Mars.
A science traverse is a planned operation: objectives, endurance, communications, return, consumables, samples and rescue capability must be prepared.

With t = 12 min one-way light-time, the minimum round trip is about 2t = 24 min. Three sequential clarifications already consume more than an hour before analysis time. Efficient Mars communication therefore bundles context, evidence, actions taken, constraints and the requested decision.

Exercise A — clarification cost

At 14 minutes one way, what is the physical minimum time for four question-and-answer cycles?

One cycle is at least 28 minutes. Four cycles require 112 minutes, or 1 h 52 min, before human reading or computation.

With many minutes of one-way delay to Mars, conversation becomes asynchronous messaging. A poorly formed question can waste an entire light-time cycle before useful analysis even starts. A message to Earth should therefore contain current state, trend, configuration, actions already taken, hypotheses and the decision being requested. Earth should respond in the same style, stating assumptions that make the advice valid. Communications time becomes an operational resource. The local team continues stabilising and measuring while remote analysis proceeds. This prevents Earth from becoming a slow remote control for a system that evolves faster than the dialogue can complete.

3. Procedure, checklist and mission intent are different tools

A procedure describes sequence; a checklist protects critical points from omission; mission intent explains what outcome must survive when the case no longer matches the script. Autonomous teams need all three. A checklist cannot invent a response to a novel failure, while broad intent can be too vague under stress.

A procedure describes a controlled method; a checklist protects against omission of critical items; mission intent explains the outcome to preserve when reality leaves the written case. Confusing these tools creates either enormous lists or uncontrolled improvisation. An irreversible action may require an observable condition and confirmation. A frequent routine may need only a short checklist. When the crew encounters an unanticipated state, intent—preserve habitable pressure before science productivity, for example—supports judgement. Training therefore teaches not only how to execute steps but how to recognise when the procedure no longer matches the actual system state.

4. Handover carries assumptions and capability debt

Shift change should transfer open anomalies, temporary configurations, consumed margins, deferred work and the rationale behind decisions. “Pump B isolated” is less useful than “Pump B isolated after unstable current; Pump A running; no redundancy until the planned 18:00 test.” The log becomes shared operational memory.

A good handover transfers more than open tasks. It carries assumptions, weak anomalies, consumed margin and deferred decisions. A pump may still operate while current slowly drifts; if that fact disappears at shift change, the next alarm will look sudden. The log should separate measured fact, interpretation and action. A practical structure can include current state, changes since the last watch, active risks, prohibited activities, decisions pending and wake-up thresholds. The goal is not to write a novel at every handover but to preserve enough operational memory that the incoming team does not have to reconstruct context under pressure.

5. Schedule real human capacity, not perfect nominal utilisation

Crew time, sleep, expertise and EVA hours are resources. A six-hour two-person repair may be impossible when another anomaly requires continuous monitoring. Plans need reserve capacity for the unplanned rather than scheduling every person to 100 percent nominal utilisation.

Human capacity can be budgeted like electrical energy. A day with 40 theoretical person-hours should not be scheduled to all 40 if the mission needs anomaly response. Reserving 20 percent leaves 8 person-hours unassigned, although the right fraction depends on phase, fatigue and risk. Presence is not the same as competence: four available people do not replace two specialists if a repair requires particular qualification. Schedules also include sleep, meals, exercise, habitat maintenance, training and documentation. A resilient mission accepts that some capacity remains deliberately unproductive during nominal operations because that reserve is what makes off-nominal work possible.

6. Automation should make safe reactions fast and ambiguous decisions explainable

Automation excels at monitoring and bounded rapid response. When judgment is required, it should expose the evidence, affected functions and confidence behind a recommendation. An alarm that says only “critical fault” adds cognitive load; an alarm that explains trend and safe options supports autonomy.

Automation is valuable when it makes a safe reaction faster and repeatable. It becomes dangerous when it hides uncertainty or performs an irreversible action from a weak diagnosis. A good interface shows what triggered the action, which evidence supports the interpretation and what condition permits return. Humans need a practical way to take over when context exceeds the model, but ‘manual control’ cannot mean hundreds of low-level commands that no tired operator can manage. Degraded modes therefore combine minimum trustworthy automation, clear procedures and a readable state representation. The objective is to reduce cognitive load rather than move opacity from one layer to another.

7. Manage anomalies as stabilise, understand, recover

Stabilise prevents further deterioration. Understand gathers evidence and narrows causes. Recover restores functions in controlled sequence. Mixing the phases can produce premature repair commands before the actual state is understood.

Anomaly management can be organised as stabilise, understand, recover. Stabilise means stop escalation even when the exact cause is unknown. Understand means collect evidence, compare hypotheses and identify dependencies. Recover returns the service to a sustainable state and proves that return. Jumping directly to repair can destroy evidence or create a second fault. The crew also preserves a timeline of commands and observations so Earth or the next watch can reconstruct the event. This sequence is particularly useful under delay because it creates a coherent anomaly package while the spacecraft remains in a controlled state.

8. Failure scenario: technical anomaly after poor sleep

A fatigued team does not have the same decision performance as a rested one. If the anomaly can be stabilised safely, the rational response may be to automate monitoring, defer complex repair, prepare tools and restore crew readiness. Human safety includes fatigue management, not just technical compliance.

Fatigue changes perception, working memory, decision speed and risk tolerance. The system should not assume the crew can always execute a complex procedure at peak cognitive performance. Critical work can be delayed when a stable mode exists; urgent actions should be short, observable and supported by clear interfaces. Cross-training also reduces dependence on the most expert individual if that person is ill or exhausted. In a combined technical failure and poor-sleep scenario, the rational first choice may be to stabilise the system and let a rested team conduct the detailed diagnosis later. Human resilience depends on the ability to buy time.

9. Ask Earth questions that remain useful after delay

Earth is valuable for deep independent analysis. The crew can ask “compare repair A and B for our present configuration” instead of “tell us what to do now.” This lets Earth exploit expertise without pretending to close a real-time loop it physically cannot close.

A question sent to Earth should still be useful when the answer returns. ‘What do we do now?’ is weak if the reply arrives forty minutes later and the state has changed. A stronger request defines branches: if temperature exceeds X, the crew will isolate the loop; otherwise it will hold the current mode until a stated time. Earth can then analyse each future state, challenge thresholds and propose actions that remain applicable. This turns delay into parallel work. Messages need units, timestamps, configuration versions and uncertainty so a recommendation is not accidentally applied to a system state that no longer exists.

10. Lessons learned need controlled change

After an incident, procedures may change. The change should retain version, rationale, evidence, approval and rollback. A successful one-off workaround is not automatically a universal rule. Operational knowledge should evolve with the discipline of a well-managed software configuration.

Lessons learned matter only when they change work in a controlled way. After an anomaly, the team separates immediate correction, probable cause, evidence obtained, procedure change and required test. A new version should identify what changed and why. Silently replacing a checklist destroys the ability to understand earlier decisions; never changing it preserves a known weakness. The process therefore combines memory with evolution. On Mars, the same principle applies to physical configuration: a maintenance procedure must know which sensor version, firmware build or locally manufactured part is actually installed before prescribing an action.

Guided case — an anomaly with 18 minutes of one-way light time

A thermal-loop pump shows increasing current but remains functional. Earth is 18 light-minutes away, so even a simple question requires at least 36 minutes before an answer returns, plus analysis time. The crew first stabilises locally: non-essential loads are reduced, the redundant loop is checked and a switchover threshold is defined. The message to Earth includes current trend, temperature, flow, configuration, recent maintenance and actions already completed.

While waiting, the local plan has branches. If temperature exceeds 65 °C, the pump will be isolated and bypass opened; otherwise the system remains under observation for one hour. Earth can analyse a future that will still be useful when the response arrives. Advice based on the assumption that the crew has done nothing would already be obsolete.

The next day, the team updates handover notes and the controlled procedure. It preserves the difference between observed fact, the hypothesis of bearing degradation and the maintenance action selected. The operational lesson becomes reusable without turning an unproven hypothesis into historical truth.

11. Mini-project: run a degraded Mars day

  1. Define four crew roles.
  2. Create a schedule with 20 percent time reserve.
  3. Add a 10:00 fault and two-hour Earth outage.
  4. Write a stabilisation checklist.
  5. Prepare the delayed Earth message.
  6. Plan evening handover.
  7. State how the procedure changes after review.

The purpose is to show that the organisation still functions when the nominal plan disappears.

12. Mission lab — build an anomaly package for delayed Earth support

A good delayed message lets an Earth team work without five clarification rounds. It includes time and configuration, initial symptom, relevant evidence, actions already taken, results of those actions, remaining resources, crew hypotheses and the decision being requested. If conditions can change before the reply arrives, the package also states which thresholds will trigger local action.

Imagine a pump whose temperature is rising. Instead of transmitting one number, the package includes several hours of trend, motor current, flow, pressure, ambient temperature, recent maintenance and the state of the redundant loop. Earth can test competing explanations while the crew retains authority for urgent thresholds already defined in local procedure.

This style turns communications into an interface between two asynchronous teams. It also makes decisions auditable because the eventual action can be traced to the evidence available at the time rather than reconstructed from memory.

Earth responses should use the same discipline. They can state assumptions, confidence, configuration for which the advice is valid and conditions that would invalidate it. That prevents a well-intended recommendation from being applied after the local state has changed.

13. Budget workload and preserve response capacity

A day scheduled to 100 percent utilisation has no resilience. If four crew members each have ten useful work hours, theoretical capacity is 40 person-hours. Reserving 20 percent for unplanned work leaves 40 × 0.20 = 8 person-hours of response capacity and 32 person-hours of nominal tasks. Twenty percent is not a universal optimum; the point is that human reserve can be budgeted explicitly rather than hoped for.

Skill reserve matters as well. If only one crew member can service a vital system, illness, fatigue or EVA can create a human single point of failure. Cross-training, clear procedures and diagnostic tools reduce that dependency. Critical maintenance should be mapped to at least the number of qualified people required by the mission’s loss-of-crew and fatigue assumptions.

Finally, protect sleep and recovery as functional resources. A tired crew can be safe if the system allows stabilise-and-wait decisions; it becomes vulnerable when every anomaly demands immediate complex manual work. Operations architecture should therefore include states in which machines hold the system stable long enough for humans to regain decision quality.

Sources and references

Engineering studio — command under latency and human workload

The scenario combines a slow habitat pressure decrease, an immobilised rover and a delayed communications pass. Two crew members are available and Earth cannot contribute a useful decision for more than 20 minutes. Workload at one station is represented by ρ = λs, where λ is task arrival rate in tasks/h and s is average task duration in hours. With λ = 4 tasks/h and s = 10/60 h, ρ ≈ 0.67: the station looks sustainable on average, yet a burst of anomalies can immediately saturate human capacity.

The task therefore goes beyond an average. The crew prioritises the leak, rover safety and communications, decides what may wait, and records authority-transfer conditions. A good procedure states what triggers volume isolation, when an EVA must be aborted and which data must be preserved for Earth. The expected result is an operational loop that remains robust when remote support is not synchronous.