The Indie Developer Who Built a 120B LLM Network Using a Phone and Two PCs
Running a 120B LLM locally normally sounds like the kind of task that requires a workstation with enormous memory, an expensive GPU, or a high-end machine with unified memory. One Reddit user decided to test exactly how far ordinary consumer hardware could be pushed instead.
The result is one of the strangest local-AI experiments to go viral this week.
Reddit user Medicine_Blogscanner, posting in the r/LocalLLaMA community, managed to run the 4-bit quantized version of OpenAI's gpt-oss-120b by distributing the model across six very different consumer devices.
The cluster included a Samsung Galaxy S24+, two Windows machines, and three Macs.
Individually, none of those systems had enough memory to comfortably hold the model. Together, they were able to run it.
There is, however, a huge catch.
The system produced only around 1.1 tokens per second, making it much more of an engineering demonstration than a practical replacement for a conventional AI workstation.
Still, the experiment demonstrates something fascinating about local AI: the memory requirement of a giant model does not necessarily have to be satisfied by one giant computer.
────────────────────────────────────────
────────────────────────────────────────
The 120B LLM Experiment
The story starts with gpt-oss-120b, OpenAI's open-weight model designed for advanced reasoning and general-purpose use.
According to OpenAI's own model repository, gpt-oss-120b contains approximately 117 billion parameters, with around 5.1 billion active parameters during inference. OpenAI describes it as a model intended to fit into a single 80GB GPU under its supported configuration.
That is already far beyond what most ordinary PCs can handle comfortably.
The Reddit experiment attacked the problem from a completely different direction.
Instead of asking:
"What single computer has enough memory?"
The experiment effectively asked:
"What if several computers share the model?"
That distinction is the entire reason the setup works.
Rather than loading the complete neural network onto one machine, the model was divided into sections and distributed across multiple devices.
────────────────────────────────────────
What Exactly Was Running?
The model at the center of the experiment was gpt-oss-120b, OpenAI's largest model in the gpt-oss open-weight family.
It is important to distinguish the model's parameter count from its actual memory requirement.
A model with 120 billion parameters does not automatically require exactly 120GB of RAM.
Quantization changes that dramatically.
The Reddit experiment used a 4-bit quantized version of gpt-oss-120b. Reporting around the experiment puts the typical memory requirement for that configuration at roughly 60GB or more, depending on the exact setup and overhead.
For comparison, an unquantized version would require dramatically more memory.
Wccftech reported that the unquantized model could require around 240GB just for the model weights under a straightforward calculation.
That is why quantization is so important for local AI.
What is quantization?
Quantization reduces the numerical precision used to represent model weights.
Instead of representing weights with larger numerical formats, a quantized model can use fewer bits.
The trade-off is that lower precision can affect performance or output quality depending on the model and quantization method.
But for enormous models, quantization can be the difference between:
"This cannot fit."
and
"Let's see if we can make this work."
────────────────────────────────────────
The Six-Device AI Cluster
The hardware involved is what makes the story so unusual.
The experiment reportedly used:
| Device | Hardware contribution |
|---|---|
| Samsung Galaxy S24+ | Android compute |
| Windows laptop | 12GB system memory |
| RTX 3060 mini PC | 12GB VRAM plus system memory |
| Intel MacBook Pro | CPU/system memory |
| Mac mini | 16GB memory |
| M3 MacBook Pro | 16GB memory |
The RTX 3060 machine acted as the primary system, while the remaining machines contributed portions of the workload.
The important point is that this was not a collection of identical servers.
It was an extremely heterogeneous collection of consumer devices.
An Android smartphone was sitting in the same distributed inference system as Apple Silicon Macs, an Intel Mac, a Windows notebook, and an RTX 3060 machine.
That is what makes the experiment so interesting.
Normally, combining completely different hardware architectures into one AI workload would sound like a nightmare.
Here, it became the solution.
Multiple reports identified the setup as a six-device cluster running the model through distributed inference and pipeline parallelism.
────────────────────────────────────────
How the Devices Worked Together
The key idea behind the experiment is pipeline parallelism.
Think about a large language model as a giant chain of processing stages.
Instead of asking one computer to perform every stage, the workload can be divided.
Machine A handles one portion.
Machine B handles the next.
Machine C handles another.
The output from one section is passed to the next device.
Eventually, the result travels through the entire model.
That is fundamentally different from simply connecting computers together and adding their RAM together.
The devices do not magically become one computer with a giant shared memory pool.
Instead, the software coordinates the workload so that different machines handle different parts of the neural network.
This is why the experiment is better described as distributed inference than simply "combining RAM."
QUICK EXPLAINER
Traditional setup:
One machine → entire model → output
Distributed setup:
Machine A → Machine B → Machine C → Machine D → Machine E → Machine F → output
The second approach introduces communication overhead, but it can allow a model to run even when no single device has enough memory to hold it.
────────────────────────────────────────
How 60GB Became Roughly 47GB
This is one of the most interesting technical details.
The experiment reportedly brought the total memory footprint down to approximately 47GB, below the roughly 60GB figure commonly associated with the 4-bit version of the model.
That does not mean the model itself somehow became a 47GB model.
The difference comes from how the workload was distributed and managed.
Not every device needed to hold the entire model.
Instead, the model's layers were divided between machines, while the system coordinated the transfer of information through the pipeline.
The primary machine reportedly carried around 28.7GB of the overall workload, with the other devices handling additional portions.
This is a crucial distinction.
The experiment did not break the laws of computer memory.
It changed the architecture of the computation.
────────────────────────────────────────
Why Pipeline Parallelism Matters
Pipeline parallelism isn't a new concept.
Large AI systems have used distributed computing techniques for years because modern models can become too large for one accelerator or server.
What's unusual here is the hardware.
The Reddit experiment effectively applied the same broad principle to machines that were never designed to operate as one AI supercomputer.
That creates several technical complications.
Different devices have different:
- CPU architectures
- Memory capacities
- Storage speeds
- Network connections
- Operating systems
- Compute performance
- Hardware accelerators
The slowest components can become bottlenecks.
Data also has to move between devices.
That introduces latency.
And because language models generate output sequentially, slow communication can become particularly painful.
The experiment proves that the architecture can work.
It does not mean that a random collection of old laptops automatically becomes a high-performance AI cluster.
────────────────────────────────────────
The Galaxy S24+ Is Part of the Cluster
Perhaps the most viral part of the story is the smartphone.
A Samsung Galaxy S24+ participated in the six-device setup.
That sounds ridiculous until you look at what the phone is actually being asked to do.
It isn't running the entire 120B model locally inside its own memory.
Instead, it contributes a portion of the distributed workload.
That distinction is critical.
The experiment therefore isn't evidence that a Galaxy S24+ can independently run a full 120B model at practical speed.
It is evidence that a smartphone can participate in a distributed inference pipeline alongside more capable computers.
That is still a remarkable demonstration.
It also hints at a broader future for local AI.
As smartphones become more capable and local inference software becomes more sophisticated, devices that normally sit idle could potentially become additional compute nodes.
Your phone doesn't necessarily need to replace your PC.
It could supplement it.
────────────────────────────────────────
The Performance Catch
Now for the part that keeps this from becoming the next big home-AI recipe.
The cluster reportedly generated approximately 1.1 tokens per second.
That is extremely slow.
For context, a typical person reading normal prose processes far more than one token per second.
So although the model technically runs, interacting with it feels nothing like using a modern cloud AI service.
You send a request.
The system processes it across the network.
The model generates output.
And the text arrives slowly.
A short response can take a significant amount of time.
A long response becomes an exercise in patience.
1.1 tokens per second is the headline
| Metric | Distributed setup |
|---|---|
| Model | gpt-oss-120b |
| Quantization | 4-bit |
| Devices | 6 |
| Approx. total memory footprint | 47GB |
| Approx. generation speed | 1.1 tokens/sec |
| Smartphone included | Yes |
| RTX 3060 included | Yes |
The speed is therefore the biggest limitation.
This is a proof of possibility, not a practical consumer workstation.
────────────────────────────────────────
Why This Is Not a Practical AI PC Yet
The viral framing makes the experiment sound like a way to get a massive AI model running on cheap hardware.
Technically, yes.
Practically, not really.
If your goal is to have a responsive local AI assistant, building a six-device cluster that produces around 1.1 tokens per second isn't an attractive solution.
The network introduces additional overhead.
The hardware has to remain powered.
The devices need to communicate reliably.
The entire system has to be configured correctly.
And the performance is nowhere near what users expect from modern AI services.
There is also another important factor.
A model running locally is not automatically better simply because it is local.
For many users, a smaller model running quickly can be far more useful than a massive model generating one token at a time.
The experiment is impressive because it answers a different question:
Can this be done?
The answer is yes.
The more practical question is:
Should you do it?
For most people, probably not.
────────────────────────────────────────
What the Experiment Actually Proves
The biggest takeaway isn't that everyone should connect their old laptops together.
It's that memory capacity is becoming an increasingly interesting software problem.
The conventional approach to local AI is straightforward:
Buy a machine with enough memory.
But distributed inference opens another possibility:
Use software to make multiple smaller machines cooperate.
This becomes particularly interesting as open-weight models continue to grow.
OpenAI's gpt-oss family is already designed around efficient inference characteristics, with only a fraction of the model's parameters active for each token because it uses a mixture-of-experts architecture. OpenAI lists approximately 5.1 billion active parameters for gpt-oss-120b despite its roughly 117-billion-parameter total size.
That architecture makes the model especially interesting for people experimenting with unconventional local inference.
And the community is already pushing the concept even further.
Independent projects such as BigMoeOnEdge demonstrate that some large mixture-of-experts models can be streamed from storage rather than requiring the complete model to remain resident in RAM. One project reports running models substantially larger than the available RAM on ordinary phones and PCs.
That means the future of local AI may not simply be about buying more RAM.
It could also involve better ways of moving data between memory, storage, accelerators, and multiple devices.
────────────────────────────────────────
Could You Build Something Similar?
Technically, enthusiasts with compatible hardware could experiment with distributed inference.
But this particular setup isn't a plug-and-play consumer product.
The original experiment required significant configuration and coordination between very different devices.
You would also need to account for:
- Network bandwidth
- Network latency
- Device compatibility
- Memory allocation
- Model quantization
- Storage speed
- Operating-system differences
- Pipeline configuration
- Thermal limits
- Power consumption
The biggest mistake would be assuming that simply putting six computers on the same Wi-Fi network automatically gives you six times the AI performance.
It doesn't.
The software architecture is doing the difficult work.
────────────────────────────────────────
The Real Innovation Is the Mindset
There is something very appealing about this experiment beyond the technical details.
Instead of looking at a 60GB-plus memory requirement and deciding that the project was impossible, the developer approached the problem differently.
The solution wasn't:
"I need a more powerful computer."
It was:
"What if I use the computers I already have?"
That kind of thinking is exactly what makes the local AI community interesting.
Open-weight models give developers access to increasingly capable systems.
Consumer hardware provides increasingly powerful CPUs, GPUs, NPUs, and unified-memory platforms.
And software developers keep finding unconventional ways to connect them.
Sometimes the resulting system is practical.
Sometimes it is painfully slow.
But experiments like this reveal what the hardware and software stack is actually capable of when conventional assumptions are removed.
────────────────────────────────────────
From a Phone to a Home AI Cluster
The most interesting part of this story may be where it leads next.
Today's experiment requires six devices to run a giant model slowly.
Tomorrow's software could make similar distributed systems considerably more efficient.
Imagine a future home setup where:
- Your desktop provides GPU compute.
- Your laptop provides additional memory.
- Your phone contributes an AI accelerator.
- An old Mac provides spare CPU resources.
- An SSD streams model data.
- Software dynamically moves workloads between devices.
That would effectively turn a collection of consumer electronics into a small personal AI cluster.
We're not there yet.
But the experiment demonstrates the basic idea.
And because open-weight AI models are increasingly available to developers, there is a growing incentive to experiment with exactly these kinds of unconventional configurations.
────────────────────────────────────────
FAQ
What 120B model was used?
The experiment used OpenAI's gpt-oss-120b, an open-weight model with approximately 117 billion total parameters.
Did one phone actually run the entire 120B model?
No. The Galaxy S24+ was one part of a six-device distributed system. It did not independently hold and run the entire model.
How many devices were used?
Six devices were reportedly involved: a Galaxy S24+, two Windows machines, and three Macs.
How much memory did the experiment use?
The reported overall memory footprint was approximately 47GB, using a 4-bit quantized configuration and distributed inference.
How fast did the model run?
The reported generation speed was approximately 1.1 tokens per second.
Does this mean anyone can run a 120B LLM on a phone?
Not independently in the same sense. The experiment shows that a phone can participate in a distributed system. It does not mean a typical phone can run the complete 120B model by itself.
Is this faster than cloud AI?
No. At roughly 1.1 tokens per second, this particular setup is primarily an engineering demonstration rather than a practical alternative to fast cloud inference.
Why use multiple computers?
The main reason is memory. No individual machine in the setup had enough resources to comfortably host the complete model, so the workload was divided across several devices.
Is gpt-oss-120b open source?
OpenAI describes the gpt-oss models as open-weight models. The weights and supporting code are publicly available through OpenAI's official repositories and model distribution channels.
────────────────────────────────────────
Final Take
A developer did not magically make a phone capable of running a 120B AI model.
What happened is arguably more interesting.
A Reddit AI enthusiast took a collection of ordinary consumer devices, including a Galaxy S24+, Windows PCs, and Macs, and turned them into a distributed inference system capable of running gpt-oss-120b.
The result is slow.
Very slow.
At around 1.1 tokens per second, nobody should mistake this for a practical replacement for a high-end AI workstation.
But that isn't really the point.
The experiment demonstrates that the boundary between "this device isn't powerful enough" and "this workload is impossible" can sometimes be crossed through software architecture.
Instead of buying one machine with enough memory, the experiment distributed the problem.
Instead of demanding one giant pool of RAM, it divided the model.
And instead of treating a smartphone as useless for a workload this large, it made the phone one piece of the puzzle.
That's the fascinating part.
The future of local AI may not simply be about building increasingly expensive computers.
It may also be about figuring out how to make all the computers we already own work together.







