Stopping GPU Dependence Means Rewriting the Stack, Not Swapping the Chip
This feature analyzes the first half of YC Paper Club's episode "What If We Stopped Using GPUs?" The episode frames GPU dependence as a product of coadaptation among hardware, Transformer workloads, memory systems, and backpropagation. It then examines zero-order optimization, the Soma sharded-expert proposal, optical matrix-vector multiplication, and Ilker's optical diffusion-generation experiment. The supported conclusion is not that GPUs are about to vanish, but that any credible alternative must redesign hardware, model structure, optimizer, and task constraints together.
1. Host and Subject Background
This is not a single-guest interview with a verified guest biography. It is an episode analysis of a YC Paper Club session published by Y Combinator on October 2, 2026, under the title "What If We Stopped Using GPUs?" The indexed evidence identifies the podcast context as Y Combinator Podcast / YC Paper Club, with Y Combinator as host and uploader. The first half of the episode is led by an opening speaker who frames the problem, then by Ilker, who presents optical computing work.
The opening speaker's role is to define the conceptual problem. He begins with a dinner-party-style thought experiment: if he could ask superintelligent aliens one question, he would ask how they compute FLOPS. That framing matters because it shifts the discussion away from buying a faster accelerator and toward the deeper question of physical computation. His central claim is that current AI sits deep inside a path created by hardware architecture and optimizer coadaptation.
Ilker's role is to turn that question into a specific substrate analysis. He is introduced as an EPFL PhD who spent about six years working on optical AI computing and was doing a postdoc at Stanford. His presentation does not claim that light is a universal replacement for electronics. Instead, it asks where optical propagation, passive weighting, high bandwidth, and low loss can be matched to an algorithm whose workload actually fits those properties.
2. What the Episode Covers
The episode's first major move is historical. The opening speaker says he has worked in deep learning since roughly 2012 and describes the CNN era as running from AlexNet until about 2019 or 2020. In his account, the earlier hardware emphasis was FLOPS and compute per joule, with crypto and Ethereum-like workloads also shaping priorities. After GPT and Transformers became central, the bottleneck shifted: attention's n-squared behavior made memory capacity and memory bandwidth much more important. He treats that shift as evidence that models and hardware incentives evolve together.
The next move is optimizer-level. The speaker says GFLOP-per-joule improvements have not advanced much in the preceding two years and uses that as motivation for alternatives to backpropagation or to the dominant compute path. He brings in the brain's roughly 20 watt power draw as a comparison point, but he does not claim that engineering should simply copy the brain. In fact, he argues against the idea that the brain performs backpropagation, citing feed-forward behavior and the weight transport problem.
He then presents zero-order optimization as a concrete alternative. In his PhD work, he says he implemented zero-order optimization methods and tested them in a one-billion-parameter LSTM next-token prediction setting. Among the methods he tried, he says SPSA performed best. The described procedure is simple in outline: perturb parameters in a random direction with positive and negative epsilon passes, use two forward passes to estimate a finite difference, and scale the update direction according to that difference.
Ilker's optical-computing presentation provides the substrate case. He says nearly all transmission between AI data center racks is optical and that 99% of intercontinental data transfer is optical. He also claims photons have about 10,000 times lower loss and 10,000 times larger bandwidth than electrons. From there, he explains the computational appeal: photons do not interact with one another, so many beams can coexist in the same space; in matrix-vector multiplication, an optical system can modulate inputs, pass them through a passive weight mask, and sum results on detectors.
The central experiment is an EPFL and Google collaboration that programs light propagation as a denoising or generation inference unit for diffusion-based image generation. The system starts from random pixels, repeatedly passes light through a fixed physical system, and trains that system to predict the noise term in noisy samples. Because optical weights are hard to reprogram frequently, the team divided 1000 diffusion steps into 10 subunits with fixed parameters reused 100 times each. Ilker says the current demonstrations are limited to small datasets such as MNIST and Fashion MNIST, but he also reports lower FID and power-law-like behavior when scaling diffractive layers or pixels per layer.
3. Core Views: Reasoning, Examples, and Limits
The load-bearing view in this part of the episode is that GPU dependence is not a single-chip problem. It is a stack problem. The opening speaker links the CNN era, the Transformer era, and the changing hardware bottlenecks to make one point: when the workload changes, the definition of a good accelerator changes with it. CNNs rewarded one set of FLOPS and energy-efficiency priorities; Transformers rewarded memory capacity and bandwidth because attention is memory-hungry and scales unfavorably with sequence length. GPU success therefore reflects a fit among backpropagation, parallel training, memory hierarchy, and model architecture.
This is why the zero-order optimization section matters. The speaker is not merely proposing SPSA as a drop-in replacement for backpropagation. He is arguing that if gradients are no longer the central training signal, then architectures designed around gradient flow may also be the wrong default. His examples support the appeal: SPSA does not require gradients, can address 0-1 loss landscapes, and in his Ackley's function example avoids the kind of local-minimum failure he attributes to backpropagation. But the episode also supplies the limitation: zero-order gradient estimates get worse as model size grows. A 10-billion-parameter model may require so many perturbations that it becomes less efficient than an ordinary forward-and-backward pass.
Soma is the proposed answer to that scaling problem. By clustering CommonCrawl with TF-IDF and SVD, sending Parquet shards to different GPUs, training smaller experts, and routing among them at test time, the speaker tries to make each optimization problem small enough that zero-order noise is capped by expert size rather than total model size. This is a reasoning move, not a finished proof. It suggests that sharding is not only an inference-efficiency trick, as in common sparse-expert systems, but potentially a way to control estimator noise during training. The uncertainty is substantial: the episode does not establish production-scale routing quality, expert coordination, global knowledge integration, or robustness under distribution shift.
Ilker's optical-computing view is similarly conditional. The strongest claim is not that light can replace all electronic compute. The stronger and better-supported claim is that light is excellent for certain physical operations: passive propagation, spatial parallelism, high-bandwidth communication, and repeated use of fixed transformations. His matrix-vector multiplication explanation gives the key intuition. Electronic systems spend energy charging and discharging wires and switching signals through defined paths; optical systems can encode inputs onto modulators, let physics apply a passive weight mask, and read summed results at detectors. That makes optical compute attractive when the algorithm can exploit fixed weights and repeated forward passes.
The limits are not side details; they are central to the argument. Ilker cites Lightmatter's optoelectronic PCIe board as showing GPU-comparable compute and power-efficiency numbers, but he then notes that photonics itself is only a small part of the energy budget. DAC conversion from digital memory, ADC conversion back to digital, calibration and programming, and nonlinear activations consume the advantage. A workload that repeatedly moves large parameter sets or large data volumes between digital electronics and optics may lose the very benefit optics promised.
The optical diffusion experiment is valuable because it is an algorithm-substrate match. Diffusion generation repeatedly applies a denoising transformation, so a fixed optical system reused across many steps is plausible. The team explicitly designed around the difficulty of rewriting optical weights by dividing 1000 steps into 10 fixed subunits. Yet the uncertainty remains large. The evidence concerns small datasets, an energy comparison that assumes passive fixed microfabricated weights, and a prototype that used a spatial light modulator and a fully digital twin for backpropagation. Ilker himself says the next milestone is billion-parameter-scale demonstration. Until that exists, the safest conclusion is that optical diffusion is a promising proof of direction, not a general replacement for GPU training.
4. Learning and Application
The practical lesson for AI builders is to stop diagnosing the compute problem as simply "not enough GPUs." The more useful breakdown is forward-pass cost, backward-pass cost, memory bandwidth, data movement, precision, and whether the task can reuse fixed weights. If the workload is large-scale training with frequent parameter updates and expensive gradient flow, optical computing as described here is not an immediate substitute. If the workload is repeated fixed inference with limited interfaces, low write frequency, and tolerable precision constraints, optical or analog systems become more relevant.
Zero-order optimization should be treated in the same conditional way. SPSA-like methods are attractive when gradients are unavailable, unhelpful, or impossible to propagate through a physical system. They can be useful for black-box objectives, nondifferentiable losses, hardware-in-the-loop systems, and small local models. They are much less attractive if every perturbation requires an expensive forward pass through a huge model. The episode's own evidence points to a boundary: zero-order training may need sharding, small experts, reproducible perturbations, or very cheap forward computation before it can compete with backpropagation.
Soma's application lesson is less about copying a named architecture and more about using data structure to reduce training difficulty. If a corpus can be clustered into meaningful shards, each expert can face a smaller local problem, and optimizer noise may become more manageable. The tradeoff is architectural complexity. A system like this would need robust clustering, a router that does not collapse or overfit, mechanisms for cross-shard knowledge, and evaluation that proves the mixture actually improves learning rather than just hiding failures in separate experts.
For optical computing, the decisive application question is interface cost. Ilker repeatedly shows that optical propagation can be efficient while the surrounding electronic system dominates energy. A serious evaluation must ask whether DAC and ADC are counted, whether memory movement is counted, whether calibration is counted, whether weight programming is counted, and whether the reported advantage depends on fixed passive weights. Without those accounting boundaries, numbers such as 10,000-times lower loss or a possible jump from seven-times to hundreds-times advantage can easily be misread as universal AI-compute facts rather than speaker-framed claims under specific assumptions.
The diffusion example gives a useful design pattern. Start with an algorithm that repeats the same transformation many times. Then ask whether fixed hardware can embody that transformation and be reused without constant reprogramming. Reduce electronic interfaces by keeping information inside the optical or analog loop longer. Accept that this narrows the task: it looks more like ASIC-style application-specific inference than general-purpose computing. That boundary is not a weakness if the workload is chosen carefully; it is the condition under which the physical advantage can survive.
The final application rule is sequence: choose the task first, choose the physical substrate second, then design the model and optimizer around the substrate's native strengths. A post-GPU strategy should ask whether the task needs frequent writes, deep nonlinearities, high-precision memory access, or general programmability. It should also ask whether a fixed-weight, low-interface, repeated-forward architecture can solve the same problem. The episode's most useful contribution is not a prediction that GPUs disappear, but a disciplined way to turn "alternative compute" into testable engineering hypotheses.
Source
- Original episode: What If We Stopped Using GPUs? | YC Paper Club
More from WayDigital
Continue through other published articles from the same publisher.
Comments
0 public responses
All visitors can read comments. Sign in to join the discussion.
Log in to comment