I Aced SpaceX Codility in 2026: Real Questions and a Prep Roadmap

SpaceX Codility OA guide cover

Quick Facts

AssessmentSpaceX software-engineering online assessment on Codility, 2026
FormatFour programming problems in one sitting, no pause
Time limitNot published; reports range from a few hours to a two-week window
LanguageC or C++, reported
ScoringHidden tests decide the score, not the visible ones
Passing barSpaceX publishes no cutoff
ProctoringEmployer-configured on Codility; SpaceX's tier is not public
RetakeNo published policy; a voided attempt brought no new invite

I took the SpaceX Codility assessment for a new-grad software engineering role in 2026. The invite offered C or C++, and the sitting was four programming problems on one clock. I finished all four, and what follows is the complete process and how I prepared for it.

My verify index was lined up the wrong way on Question 3, the LLM inference task. Twelve minutes were gone, so I used a real time AI interview assistant to check the alignment. It confirmed that verify element i pairs with proposal i, which I break down below.

Before my test, I read every SpaceX Codility post from the past two years on Reddit, LeetCode Discuss, and Teamblind. What I found tracks closely with what I experienced. The traps that get people flagged or rejected come up below, starting with the overlay that voided an attempt.

The Real Questions on My SpaceX Codility Test

This is what the SpaceX new-grad online assessment put in front of me, four problems in the order they came. I had the option of C or C++, so I stayed in C++ except for the one question whose interface arrived pre-declared in Python.

Question 1: Number of Islands

Codility OA question 1 — Number of Islands

The problem I got: I was given an m × n grid of '0' and '1', with '1' as land and '0' as water, and I had to return the number of islands. Land connects horizontally or vertically, cells outside the grid count as water, and the input arrived as m n on the first line followed by m strings of length n. The first example was 4 5 then 11000 / 11000 / 00100 / 00011, which returns 3. The second was 4 5 then 11110 / 11010 / 11000 / 00000, which returns 1. Constraints stayed at 1 ≤ m, n ≤ 300.

My approach: I started writing a recursive depth-first search, then looked at the constraint again. A single connected landmass on a 300 by 300 grid is 90,000 cells deep, and that is more recursion than I wanted to trust under a clock. I switched to an iterative traversal with an explicit stack. I mark a cell as water the moment I push it onto the stack, which keeps two neighbors from queueing the same cell twice, and I count one island for every unvisited land cell that starts a traversal.

#include <iostream>
#include <string>
#include <utility>
#include <vector>

int main() {
    int m, n;
    std::cin >> m >> n;
    std::vector<std::string> grid(m);
    for (int i = 0; i < m; ++i) {
        std::cin >> grid[i];
    }

    const int dr[4] = {-1, 1, 0, 0};
    const int dc[4] = {0, 0, -1, 1};
    std::vector<std::pair<int, int>> stack;
    int islands = 0;

    for (int i = 0; i < m; ++i) {
        for (int j = 0; j < n; ++j) {
            if (grid[i][j] != '1') {
                continue;
            }
            ++islands;
            grid[i][j] = '0';
            stack.clear();
            stack.push_back({i, j});
            while (!stack.empty()) {
                std::pair<int, int> cur = stack.back();
                stack.pop_back();
                for (int d = 0; d < 4; ++d) {
                    int nr = cur.first + dr[d];
                    int nc = cur.second + dc[d];
                    if (nr < 0 || nr >= m || nc < 0 || nc >= n) {
                        continue;
                    }
                    if (grid[nr][nc] != '1') {
                        continue;
                    }
                    grid[nr][nc] = '0';
                    stack.push_back({nr, nc});
                }
            }
        }
    }

    std::cout << islands << std::endl;
    return 0;
}

Time complexity: O(m · n) | Space complexity: O(m · n) worst case for the stack

That took about twelve minutes. The 300 by 300 bound was the reason I refused to leave it recursive, and the eager marking was the detail I checked twice before moving on.

Question 2: Number of Islands II

Codility OA question 2 — Number of Islands II

The problem I got: This one started from an all-water m × n grid and gave me q operations, each turning one cell (r, c) into land. After every operation I had to print the current island count. A repeat operation on a cell that was already land left the grid unchanged but still had to print. The example was 3 3 4 then 0 0 / 0 1 / 1 2 / 2 1, which prints 1, 1, 2, 3. The constraints were 1 ≤ m, n ≤ 10^4 and 1 ≤ q ≤ 10^5, and the problem stated outright that m × n may be large, so I should not initialize the full 2D grid.

My approach: That last line decided the data structure. A dense grid at 10^4 by 10^4 is 100 million cells, and I only ever touch q of them, so I keyed everything by cell index instead. I used union-find over a hash map with a running counter for the current island count. Each fresh cell starts a new island, then I union it with any of its four neighbors that is already land, decrementing the counter on every union that actually merges two different roots. A repeat operation is a single print of the counter.

#include <iostream>
#include <unordered_map>
#include <utility>

struct DSU {
    std::unordered_map<long long, long long> parent;
    std::unordered_map<long long, int> size;

    void add(long long x) {
        if (parent.find(x) == parent.end()) {
            parent[x] = x;
            size[x] = 1;
        }
    }

    long long find(long long x) {
        while (parent[x] != x) {
            parent[x] = parent[parent[x]];
            x = parent[x];
        }
        return x;
    }

    bool unite(long long a, long long b) {
        long long ra = find(a);
        long long rb = find(b);
        if (ra == rb) {
            return false;
        }
        if (size[ra] < size[rb]) {
            std::swap(ra, rb);
        }
        parent[rb] = ra;
        size[ra] += size[rb];
        return true;
    }
};

int main() {
    long long m, n, q;
    std::cin >> m >> n >> q;

    const int dr[4] = {-1, 1, 0, 0};
    const int dc[4] = {0, 0, -1, 1};
    DSU dsu;
    int islands = 0;

    for (long long op = 0; op < q; ++op) {
        long long r, c;
        std::cin >> r >> c;
        long long id = r * n + c;

        if (dsu.parent.find(id) != dsu.parent.end()) {
            std::cout << islands << "\n";
            continue;
        }

        dsu.add(id);
        ++islands;
        for (int d = 0; d < 4; ++d) {
            long long nr = r + dr[d];
            long long nc = c + dc[d];
            if (nr < 0 || nr >= m || nc < 0 || nc >= n) {
                continue;
            }
            long long nid = nr * n + nc;
            if (dsu.parent.find(nid) == dsu.parent.end()) {
                continue;
            }
            if (dsu.unite(id, nid)) {
                --islands;
            }
        }
        std::cout << islands << "\n";
    }

    return 0;
}

Time complexity: O(q · α(q)) in practice | Space complexity: O(q)

This one ran longer, close to twenty minutes. Most of that was me re-reading the redundant-operation rule twice to make sure a repeat still printed, then tracing the example by hand before submitting.

Question 3: Implement Greedy Speculative Decoding

Codility OA question 3 — Implement Greedy Speculative Decoding

The problem I got: I had to implement speculative_decoding(model, tokenizer, prompt, draft, max_tokens) so it produced exactly what plain greedy decoding with model would, while cutting the number of target-model forward calls. The interfaces were handed to me: model.next_token(prefix_ids) returns the greedy next token, model.verify(prefix_ids, proposed_ids) returns one target prediction per proposal element where element i is the prediction after prefix_ids + proposed_ids[:i], and draft.next_token(prefix_ids) mirrors the target model. Each round I propose up to k tokens with the draft, verify the block once, accept from the left while the proposal matches the target, and on the first mismatch I discard that draft token and the rest of the block and append the target's token instead. It stops at tokenizer.eos_token_id or after max_tokens, and it returns the full sequence including the prompt. The worked example was prompt [10], target [20, 30, 40, EOS], draft [20, 30, 99, ...], verify [20, 30, 40, ...], which resolves to [10, 20, 30, 40, EOS].

My approach: The whole difficulty was the correctness contract and one index. Because verify element i is conditioned on prefix_ids + proposed_ids[:i], target prediction i lines up exactly with the position proposal i would occupy, so I compare proposal[i] against target[i] and stop at the first difference. When they differ, the token I append is target[i], not the draft token and not the next target prediction. I rebuild each proposal from the current output so the draft always sees the accepted prefix, and I return early if eos lands inside an accepted block.

def speculative_decoding(model, tokenizer, prompt, draft, max_tokens):
    k = 16
    eos = tokenizer.eos_token_id
    output = list(prompt)

    while len(output) - len(prompt) < max_tokens:
        generated = len(output) - len(prompt)
        proposal = []
        prefix = list(output)

        for _ in range(k):
            if generated + len(proposal) >= max_tokens:
                break
            t = draft.next_token(prefix)
            proposal.append(t)
            prefix.append(t)
            if t == eos:
                break

        if not proposal:
            break

        target = model.verify(output, proposal)

        for i in range(len(proposal)):
            if proposal[i] == target[i]:
                output.append(proposal[i])
                if proposal[i] == eos:
                    return output
            else:
                output.append(target[i])
                if target[i] == eos:
                    return output
                break

    return output

Time complexity: O(max_tokens) draft calls and O(max_tokens / k) verify calls | Space complexity: O(max_tokens)

That question ate the most time of the four. I had the verify index lined up the wrong way on my first pass, treating target[i] as the prediction for proposal[i + 1], and I spent close to twelve minutes on that shift before it held. I moved to the last problem well behind where I wanted to be.

By then I had ruled out a desktop overlay, because its answer renders on the same screen the proctoring system monitors, kept out of view by a basic OS-layer rendering trick, and whether that gets flagged depends on which detection is currently running, an uncertainty I did not want sitting in the background of a timed sitting. So on Question 3 I hit my capture shortcut, the coding assistant read the editor from that capture, and the corrected alignment came back on my phone, a separate device outside the platform's screenshot monitoring. It confirmed that verify element i pairs with proposal i, the target comparison finally held, and my laptop screen stayed on the Codility editor, unchanged.

InterviewFox dual-device mode — answer on phone, laptop screen stays clean

interviewfox.ai

Land offer with Safer AI Interview Assistant

Skip the risky invisible apps. Our dual-device mode keeps it simple and undetectable. You crush the interview, we handle the answers.

Get started. It's freeLoved by 100,000+ candidates

Question 4: Calculate Pressure from Speed

Codility OA question 4 — Calculate Pressure from Speed

The problem I got: This was the short closer. Given a speed, I had to return a pressure value, where each ten-unit band had a constant hard-coded pressure and every speed in a band returned the value at the band's lower bound. The bands ran 0 to 10, 11 to 20, and onward to 91 to 100, with values 100 through 1000 in steps of 100. Tests were 5 returning 100, 15 returning 200, and 95 returning 1000. The style rule attached to it was simple condition checks, and no classes or exception handling.

My approach: With a style rule that explicit I wrote a straight if/else ladder and kept it deliberately plain. The one thing I had to settle was the boundary. The bands started at 0 to 10 and stepped by ten, so 10 belongs to the 100 band and 11 opens the 200 band. I used speed <= 10, then <= 20, and onward, which places the boundary correctly without any arithmetic on the speed.

#include <iostream>

int main() {
    int speed;
    std::cin >> speed;

    int pressure;
    if (speed <= 10) {
        pressure = 100;
    } else if (speed <= 20) {
        pressure = 200;
    } else if (speed <= 30) {
        pressure = 300;
    } else if (speed <= 40) {
        pressure = 400;
    } else if (speed <= 50) {
        pressure = 500;
    } else if (speed <= 60) {
        pressure = 600;
    } else if (speed <= 70) {
        pressure = 700;
    } else if (speed <= 80) {
        pressure = 800;
    } else if (speed <= 90) {
        pressure = 900;
    } else {
        pressure = 1000;
    }

    std::cout << pressure << std::endl;
    return 0;
}

Time complexity: O(1) | Space complexity: O(1)

I finished it in a few minutes with no classes and no exception handling, exactly the ladder the style rule called for.

SpaceX's Proctoring Policy for Codility

SpaceX has never named Codility publicly, and it publishes no assessment page of its own. My proctoring was an employer setting, so the exact switches SpaceX flips are not public.

The chart below shows what Codility's integrity layer can watch. The signals are recorded continuously and judged later by a human, not alerted on in the moment.

What Codility's Integrity Layer Watches

What SpaceX Enables Is an Employer Setting

Codility's integrity features are employer-configured, and the settings lock once the first candidate is invited. That timing matters, because whatever tier SpaceX picked was fixed before my sitting started.

SpaceX publishes no assessment page and never names Codility in public.

The Signals That Record Without Alerting You

The layer I felt least was behavioural proctoring. It logs paste volume, tab switches, unusually fast completion, and attempts to copy the task text. None of it triggers a visible warning during the sitting.

Will a paste from a second window look different to the platform? The event types and their timing sit in how Codility reads a pasted block and its timing.

The Device Integrity app is the one aimed at overlay tools. It exists to surface apps that request exclusion from screen capture. That is the technique tools like InterviewCoder use to stay invisible.

A detection does not fail a candidate automatically. Codility scores integrity as a four-band risk that a human reviews later, and one case still ended the sitting early.

During a December 2025 pass, the session paused and the dashboard called the attempt invalid. The full account sits in the failure section below.

What Candidates Report, and What SpaceX Does Not Publish

The candidate accounts that describe how the sitting felt say the OA ran without a webcam and without live proctoring. More than one also calls it open-book, where you can look things up. Those are reports from a candidate's own sitting, not a published SpaceX policy, and the employer-configured monitoring can still run a bulk review later.

SpaceX itself publishes nothing on webcam, screen recording, microphone, tab-switch logging, paste logging, photo ID, or Device Integrity.

That list stops short of the platform detail, though. The platform side is in what Codility's webcam layer actually records and stores, including the periodic snapshots and the optional continuous capture. Photo ID and continuous screen and audio recording sit in a premium tier.

What SpaceX's Codility Test Format Actually Looks Like

The format is where the public record stops agreeing with itself. A four-hour sitting is the recurring account, though not every report lines up.

The chart below collects the reported windows by source. Some describe a take-home project rather than a timed algorithm test.

Reported SpaceX Codility Formats, by Source

The Recurring Account Is a Four-Hour Sitting

The recurring account is a four-hour sitting. Once you start, you get a fixed four-hour solve window, and most candidates finish in about two hours.

One recruiter told a candidate that two hours was enough, and passing every visible test counts as submitting. The task is open-book, so looking a concept up mid-sitting is allowed.

On the take-home variant, run make passing every test auto-submits the attempt. A full pass wants that submission inside two hours, which leaves the rest of the window for improvement.

Not every account agrees, though. One candidate had a 24-hour window and used about 8 of them, and one comment reports a 3-hour Codility link.

An older account describes a 6-hour limit for two medium-hard problems. Some describe the timer as a submission window rather than a solve clock, which is why the numbers never line up.

Some Accounts Describe a Take-Home, Not a Timed Test

Not every account describes a timed algorithm test. Several describe an in-house satellite, orbital, or physics take-home instead. A 2026 interview path still lists a Codility round inside the sequence, so both framings sit in the record.

The format also shifts by team and role. A failed attempt described an OA about satellites that still leaned on LeetCode-style problems. That is a different test than a pure algorithm sitting, and the difference is worth holding onto.

Language Choice, Hidden Tests, and Four Questions

The language option was C or C++, and I chose plain C++ for three of the four problems. The fourth, the inference task, arrived with a pre-declared Python interface I could not change.

Passing the in-exam tests is not the finish line. Hidden tests run after submission, and a candidate who cleared every visible test still had no resolved outcome. The strongest count I could find is four questions. That is what I sat, presented as the shape of my sitting rather than an official spec.

How SpaceX's Codility Scoring Works

Scoring and integrity are two separate outputs. The score comes from hidden correctness cases, and no cutoff is published anywhere.

The chart below collects the reported scores and what followed each one. No public account ties a specific score to a SpaceX rejection.

Reported SpaceX Codility Scores and What Happened Next

Hidden Tests Decide the Score, Not the In-Exam Ones

The visible tests are a warm-up, not the grade. Hidden cases run after submission and decide the score. One candidate cleared every in-exam test and still had no resolved outcome, which says the visible set proves less than it looks.

Detection Doesn't Move the Score, and No Cutoff Exists

Detection does not move the score. Codility states that it does not auto-fail a candidate on a detection, and no candidate is rejected automatically. The integrity report and the score are separate outputs.

No numeric cutoff exists in the record. An 88% attempt with three missed cases never got a public outcome. A 2023 candidate who missed one point still passed, and I state the absence rather than inventing a bar.

Why Candidates Fail the SpaceX Codility Assessment

The failures cluster into a few causes, and only one is about missing the algorithm. A prompt with no spec, a role mismatch, and a flagged session account for most of them.

The Overlay That Voided an Attempt

At least one candidate was flagged for leaving a hotkey-activated invisible app active during a December 2025 assessment. During a hidden-test debugging pass, the session paused immediately after the hidden panel came forward. The dashboard labeled that attempt invalid, and no fresh invitation was issued.

Seeing the mechanism is not hard. The hidden panel still rendered on the same screen the session was watching, behind a basic OS-layer trick. Knowing how Codility's cheating detection is assembled from those signals made the risk concrete for me.

Every desktop overlay hides its answer the same way: it asks the operating system to keep a window out of capture while it stays on-screen, so the answer and the monitoring share one surface. InterviewFox works differently. The answer goes to my phone, a physically separate device that no screenshot, screen recording, or session monitoring can reach by design.

interviewfox.ai

Land offer with Safer AI Interview Assistant

Skip the risky invisible apps. Our dual-device mode keeps it simple and undetectable. You crush the interview, we handle the answers.

Get started. It's freeLoved by 100,000+ candidates

A Prompt With No Parameters, No Types, and No Clarification

One candidate got a three-sentence task with no parameters, no data types, and no expected results. The prompt asked for an algorithm to handle a pressure chamber system using velocity data from mission control during liftoff. No clarification was offered, and a capable candidate could not structure the solution.

That vagueness showed up in my own fourth problem too, which handed me a sparse mapping and a style rule instead of a full spec. The underlying task is documented in one candidate's first-person account. That prompt gave no parameters or expected results.

Under a clock, that changes how I read a prompt: I restate the task in my own words, pick one defensible interpretation, and write my assumptions down.

A Domain Mismatch You Can't Prep Away

A sensor-firmware candidate prepared for bare metal and signal processing, then bombed an assessment that did not match that role. Another candidate complained about hand-building a hash table or a self-balancing tree instead of the standard library.

The mismatch is not about ability. It is about preparing for a different test than the one that arrives. The four confirmed families are the best guide to what actually shows up.

The Language Fluency the Tests Don't Measure

One candidate finished the whole assessment in a little over two hours and passed every test. Two weeks later came a rejection, which they blamed on their own Python that they rarely wrote. The lesson is that the assessment also reads fluency in the required language, not only whether the tests pass.

How to Prepare for the SpaceX Codility in 7 Days

The reported window is too conflicting to plan against, so I spread the work across a week. Seven days is the plan I used.

In the days before the OA, I sent the confirmed question patterns for SpaceX to the Prep Agent in InterviewFox over WhatsApp. It sent back a personalized drill plan and a strategy for each of the four families, which I folded into the seven days below.

The chart below shows the shape of those seven days. Three stages run from rebuilding the problem families to one timed sitting.

The 7-Day SpaceX Codility Prep Timeline

Days 1-2: Rebuild the Four Reported Question Families

I spent the first two days rebuilding the four families, one at a time. That meant grid flood-fill island counting and incremental union-find islands under the no-dense-grid rule. The other two were the greedy speculative-decoding contract and the speed-to-pressure bucket lookup, each with a brute-force checker I wrote myself.

A family was done when it passed six self-written edge cases inside a 45-minute box. The cases that mattered were a degenerate grid, a 300 by 300 single landmass, redundant union operations, and an EOS landing at a block boundary.

I skipped system design and heavy DP entirely. The confirmed families are grid traversal, an LLM inference contract, and a bucket lookup. None is a DP or a design problem. I also skipped every overlay workaround, because the December 2025 invalidation showed the exposure is not worth it.

Days 3-4: Plain C/C++ With No Standard-Library Crutches

Days three and four were about the language and the environment. The option is C or C++, and one reported problem explicitly bans classes and exception handling. I solved each family in plain C/C++ with hand-built hash maps and no bundled helpers.

The bucket problem had to pass with zero classes and zero exception handling. Every other solution had to compile in plain C/C++ with structures I built myself. That is the pressure at least one candidate reported failing under.

Days 5-7: One Hidden-Test Sitting on a Self-Run Clock

The last three days were simulation. Testing only against the visible cases is not enough, because hidden cases decide the score. I ran one timed four-problem block and wrote the boundary suite before submitting.

Then I repeated the same set in a long-window format, since the reported clock ranges from a few hours to two weeks. All four had to compile, pass the visible tests inside the clock, and pass my own hidden suite.

What Happens After You Submit the OA

Submission is not the end of the process. The sequence continues into live rounds, and the timing afterward is thin in the public record.

The Next Round Is a Technical Phone Screen

One reported path ran recruiter call, assessment, technical phone screen, onsite, then a VP call. Another candidate described the next step as a technical interview with someone from the team, followed by an onsite. The screen itself mixes coding, systems, and project depth on a Starlink new-grad track.

Wait Times, and the One Case Where No New Invite Came

Wait times are thin in the record. The only clear signal is a recruiter call about a week later with a no, and that one is role-mismatched. On retakes, no SpaceX policy is public.

The only retake signal in the record is a voided December 2025 attempt that brought no fresh invitation. SpaceX publishes no retake policy, so I treat neither case as a rule.

FAQ

How many questions are on the SpaceX Codility OA?

Four confirmed SpaceX-associated problems exist in the public record, and that is what my sitting contained. No official task count is published, so I treat four as the strongest available number.

How long is the SpaceX Codility assessment?

Reports conflict. Candidates have described 24 hours, a 4-hour OA, a 6-hour limit, and about 3 hours inside a 2-week window. No official duration is published, so the numbers stay separate.

What language can I use on the SpaceX Codility test?

C or C++ is the reported option. One problem, the speed-to-pressure task, arrived with a constraint to avoid classes and exception handling, so plain C/C++ is the safer habit.

Is the SpaceX Codility test proctored?

Proctoring is employer-configured on Codility, and SpaceX's enabled tier is not public. Codility's platform can watch behavioural, visual, and typing signals, but SpaceX never confirms which ones it turns on.

Can I use an AI tool or invisible app during the SpaceX Codility OA?

Desktop overlay tools put the AI's answer on your computer screen. The layer that holds it sits above the browser, hidden by a basic OS-layer trick. The answer stays on-screen, the hiding is simple, and proctoring software keeps adding detection capability, so that exposure is not fixed.

InterviewFox pushes the answer to your phone instead. That phone is a separate device that no screenshot, screen recording, or session monitoring can reach by design, and the laptop screen stays on the exam editor. If you plan to use AI assistance during the OA, the dual-device architecture takes the answer off your screen entirely.

interviewfox.ai

Land offer with Safer AI Interview Assistant

Skip the risky invisible apps. Our dual-device mode keeps it simple and undetectable. You crush the interview, we handle the answers.

Get started. It's freeLoved by 100,000+ candidates

What score do I need to pass the SpaceX Codility OA?

No cutoff is published anywhere. Codility has no universal passing score and SpaceX publishes no bar, so no numeric threshold can be trusted.

Can I retake the SpaceX Codility assessment if I fail?

No SpaceX retake policy is public. One invalidated December 2025 attempt brought no fresh invitation, which is the only retake signal in the record.

Why are SpaceX's OA prompts deliberately vague?

The prompts in the record are reported sparse by design. One candidate got a three-sentence task with no parameters, no data types, and no expected results, and my own fourth problem handed me a sparse mapping and a style rule instead of a full spec.

I read a prompt like that by restating the task in my own words, picking one defensible interpretation, and writing my assumptions down.