Modern Drug Development Through Bimagrumab: Why Biology Is Only the Beginning
Peter Attia’s interview with Lloyd Klipstein uses Bema/bimagrumab as a working case study to explain how modern drugs move from unmet medical need to target biology, antibody discovery, IND-enabling work, GMP manufacturing, early clinical trials, obesity combination therapy, preventive medicine, and cancer-prevention hypotheses. The episode’s main lesson is not that one molecule is magical, but that drug development is a chain of multiplied scientific, measurement, manufacturing, regulatory, ethical, and commercial risks.
1. Guest Background
This episode of The Peter Attia Drive, hosted by Peter Attia, is titled “409 ‒ Inside modern drug development: the science, economics, and regulatory hurdles.” It is not a narrow conversation about a single obesity or muscle drug. It is an episode analysis of how modern medicines are conceived, tested, manufactured, regulated, protected, and sometimes commercialized after passing through many ways to fail.
The guest is Dr. Lloyd Klipstein, whom Attia introduces as a physician scientist, rheumatologist, drug developer, and current CEO of Coslap Therapeutics. His career spans academic medicine at Harvard Medical School and Brigham and Women’s Hospital, translational medicine and new indication discovery at Novartis, leadership roles in biotechnology companies, and development work on Bema at Versantis Bio before the company was acquired by Eli Lilly.
That background matters because Klipstein is not speaking only as a target biologist or only as a biotech executive. He has seen the full chain from identifying an unmet medical need to testing a molecule in humans and thinking about regulatory, manufacturing, investment, and commercial constraints. Attia frames Bema/bimagrumab as the recurring case study, but the broader purpose is to show why drug development takes so long, costs so much, and fails so often even when the underlying biology is real.
2. What the Episode Covers
Klipstein begins by moving the origin of drug discovery away from the molecule and toward the patient. He says his own process usually starts with patients, clinical indications, and unmet medical need: the developer should be asking what therapy does not yet exist but is needed. That distinction shapes the whole interview. Incremental innovation, such as less frequent dosing, an oral version of an injectable therapy, or a better member of an existing class, is easier for the industry to support because the disease category, endpoint, clinician behavior, and commercial pathway are already familiar. The harder work is what Klipstein calls a larger conceptual step: recognizing an indication that may not yet have a clear code, a shared clinical language, or a regulatory path.
At Novartis, his new indication discovery unit made that logic explicit. The team listed roughly 7,000 clinical indications, grouped them into categories, and looked for opportunities that the company was not addressing. Healthy aging and muscle-related disorders emerged from that process. Klipstein gives frailty as a motivating example: at the time he was examining it, he says frail older adults who had to enter a nursing home had a three-year mortality rate approaching 90%, which he describes in the episode as worse than cancer. That number should be treated as a speaker-cited framing claim from the episode, not as an independently established universal statistic, but it explains why he believed frailty deserved more serious drug-development attention.
The episode then maps the basic terrain of drug types. Klipstein explains that small molecules are essentially chemicals; biologics include antibodies, peptides, soluble receptors, and other non-chemical therapeutic formats; gene therapies are becoming their own complex category; and devices are regulated separately through FDA’s CDRH. This taxonomy is not cosmetic. The therapeutic format affects target choice, manufacturability, dosing, regulatory review, intellectual property, toxicology, and the kinds of risks developers must reduce before treating people.
The bimagrumab story begins, surprisingly, with a failed fall-prevention program. Klipstein first wanted to prevent falls in frail older adults, but falls are a messy endpoint. Weakness, dizziness, vision, attention, proprioception, reaction speed, and foot speed can all contribute. The team tried to make falls objectively measurable with a triaxial accelerometer. In a study of 60 high-risk older adults followed for six months, about 117 events were captured based on people being found on the floor. The device detected only about 17% of real falls and produced unreliable alerts. Because the endpoint could not be measured well enough, the fall-prevention drug program stopped and the work shifted toward sarcopenia, muscle mass, and strength.
Klipstein describes sarcopenia as an evolving field, with definitions moving toward decreased muscle mass plus impaired muscle function. Function might be assessed by grip strength, gait speed, stair climbing, or related measures. The target biology came from the myostatin and activin system. Myostatin inhibits muscle growth, and blocking it can make animals muscular. In humans, however, Klipstein says the biology is more complex: myostatin is not the whole story, and activins, especially activin A, matter. The development team therefore chose to block the activin type II receptor rather than target only a single ligand. The receptor system belongs to the TGF-beta superfamily and signals largely through SMAD transcription factors that regulate muscle-related gene programs, protein synthesis, and protein turnover.
The choice of an antibody followed from the binding problem. Myostatin and activins bind their receptors at nanomolar, low-nanomolar, or high-picomolar affinities, so an inhibitor has to bind even more tightly to block them effectively. Klipstein says that would be difficult with a small molecule, making a therapeutic antibody the more plausible format. The discovery project used phage display and began with thousands of candidate antibodies. Because activin type II receptors were expressed at levels too low to measure easily by flow cytometry, the team built a firefly luciferase reporter assay: cells glowed when myostatin or activin signaling occurred, and candidate antibodies were tested for their ability to reduce that signal.
Even a fully human antibody does not eliminate immune uncertainty. Klipstein explains that antibodies generated in vitro can still contain sequences the human immune system recognizes as foreign. Developers can use in silico methods to estimate risks such as peptide binding to MHC molecules, but the real immunogenicity risk is not known until human studies. Before that point, a program also needs IND-enabling toxicology and DMPK work to show where the drug goes, how it behaves, whether it reaches the target, and what organ-level toxicity appears in animals.
Manufacturing is part of that regulatory package, not an afterthought. Klipstein explains that an IND requires manufacturing information, and GMP requires documentation of identity, label accuracy, purity, activity, absence of contamination, sterility, consistency with clinically tested material, and a controlled supply chain. A typical drug may involve separate companies for drug substance, formulation, packaging, and distribution. The product is not merely the nominal molecule; it is the molecule plus a documented system showing that patients receive what was tested.
That leads Attia and Klipstein into a warning about gray-market peptides. They distinguish regulated GMP medicines from online products marketed as retatrutide or other research peptides. Buyers cannot know from the label whether the vial contains the stated compound, the intended dose, or contaminants. They also discuss BPC-157 as a wellness-market example with major evidence problems: Attia says the data are concentrated in one investigator, the findings have not been independently reproduced, the peptide is not encoded in the human genome, and it has no known receptor. The practical lesson is that “same molecule” claims are weak when identity, purity, dose, sterility, and evidence cannot be audited.
The animal results for bimagrumab were striking. Klipstein says related antibody constructs caused substantial muscle hypertrophy in mice, possibly more than pure myostatin knockout, with mice becoming stronger and faster and muscle mass increasing by roughly 30%. But humans were not mice. In tested human populations, mostly older adults, muscle mass generally increased by about 4%-8%. Novartis had strong confidence in the first-in-class biology and ran about 16 phase-one/two or phase-two studies across indications. The result was consistent: bimagrumab reliably increased muscle size, but performance measures did not improve in a major way. A meta-analysis across Novartis studies found that a 4%-8% muscle increase in older adults corresponded to a nine-meter improvement in six-minute walk distance, which both Attia and Klipstein doubted would be enough to prevent falls.
Safety and ethics added another layer. In animal toxicology, rats treated with bimagrumab had larger hearts, but they also had larger bodies and skeletal muscles. When heart size was normalized to body size, it appeared normal, and both muscle and heart size decreased after drug withdrawal. The team argued to regulators that this was not specific cardiac toxicity. More broadly, Klipstein says toxicology decisions use weight of evidence and risk-benefit judgment, with special attention to whether toxicities are monitorable and reversible. Irreversible cardiac or neurologic toxicity would usually end a program.
For first-in-human studies, participant choice depends on risk. Klipstein prefers the cleanest possible population when risk is very low, because measurements are easier to attribute to the drug rather than disease or confounding medicines. But he says he would not expose healthy volunteers to serious risk above roughly the annual U.S. lightning-strike risk of one in 100,000. If the risk is higher, the study should move into patients who might benefit. The TeGenero/CD28 agonist antibody incident, where healthy volunteers experienced severe immune activation, is cited as a reason modern first-in-human studies use sentinel participants and staggered dosing cohorts.
Bimagrumab’s phase-one work used older healthy volunteers and measured muscle mass, strength, DEXA or MRI, and soluble muscle proteins such as CK, aldolase, and LDH. Reproducible on-target adverse effects included muscle spasms or cramps, acne, and first-dose-related diarrhea. Later work suggested that nutrition and behavior could strongly modulate the drug’s effect. In a nutrition study referenced by the Believe obesity paper but not yet peer reviewed in full, more protein intake within tested boundaries produced more muscle gain, while low protein and calorie intake caused expected muscle loss that bimagrumab prevented.
The program then found a new metabolic direction. In a Novartis study of people with type 2 diabetes, bimagrumab was given at 10 mg/kg monthly for 48 weeks. Klipstein says the study showed expected muscle gain, substantial fat-mass reduction, and an absolute HbA1c decrease of about 0.7%-0.8%. Novartis later spun out the asset, and Versanis licensed it around 2021. The original plan was to treat older adults with low muscle mass, impaired muscle function, and obesity. But before semaglutide changed obesity care, investors viewed obesity drug development as commercially barren. Klipstein says the team spoke with about 53 investors and raised about $70 million, short of the roughly $100 million they wanted.
Novo’s semaglutide data changed the strategy. Versanis concluded that incretin agonists would become standard obesity treatment, so bimagrumab had to be positioned on top of GLP-1 or related therapy, with the goal of increasing fat loss and preserving lean mass. Mouse studies suggested additive effects with semaglutide, tirzepatide, and liraglutide. The Believe study expanded from an initial 24-week monotherapy concept into a 72-week treatment study with a 48-week primary endpoint, a six-month off-drug follow-up, about 507 participants, and a nine-arm full factorial design testing low/high semaglutide, low/high bimagrumab, and placebo combinations.
Klipstein personally wanted waist circumference as the primary endpoint because he considered it more closely linked to important clinical outcomes than BMI or body weight. The study used body weight anyway because regulators, investors, and potential acquirers cared about the approval endpoint. In the high-dose semaglutide plus high-dose bimagrumab group, body weight fell about 22%-23% at 72 weeks, and fat loss reached 45.7% of starting body fat, which Klipstein compares to bariatric-surgery-level fat loss. Believe also measured function and selected grip strength because it was closely linked to clinical outcomes, but the strength improvement was small and variable. The major unexpected adverse signal was an LDL increase of about 20%, which Klipstein believes is an on-target liver effect, although he says the mechanism is not yet known. He says that if Versanis had remained independent, he would have advanced the program into phase three because LDL can be monitored and treated. Eli Lilly acquired the asset, and public clinicaltrials.gov information still shows a complex combination study with tirzepatide.
The interview ends by widening from obesity to prevention. Klipstein imagines future obesity care as induction and maintenance of remission: combination and injectable therapies could move a person from obesity to non-obesity, followed by an oral GLP-1 drug or a lower-dose injectable to maintain appetite and satiety control. On mTOR, he thinks mTORC1 inhibition probably has geroprotective potential in humans, though likely with modest effect size. The hard problem is true mTORC1 selectivity without sustained mTORC2 downregulation. He also notes that young rodents downregulate mTOR with fasting whereas old rodents did not in the work he discusses, making him question whether intermittent fasting produces the same biology in older people.
Klipstein’s current cancer-prevention company is built on a reverse inference: if some drugs cause cancer, they may be inhibiting cancer-protective pathways. He cites sorafenib and related multikinase inhibitors causing skin cancers in about 10% of older patients as the clue. One inhibited target involves a sensing kinase that triggers ribotoxic stress, a cell-death pathway. His hypothesis is that gently and controllably increasing that pathway could prevent some cancers. He explicitly frames a possible 50% prevention hope as a hypothesis, not a proven result. A proposed phase-two study would start with adults aged 50 or older who have had at least five prior skin cancers, because he says they have about a 50% chance of another skin cancer within one year and could be studied in cohorts of roughly 100-120 people. A collaboration with the Broad Institute tested a tool compound across about 1,000 cancer cell lines and found melanoma to be the most sensitive tumor type, although the mechanism is not fully explained. Attia and Klipstein also stress that conventional cancer prevention still includes avoiding carcinogens and improving metabolic health: not smoking, avoiding alcohol, maintaining insulin sensitivity, and losing weight when overweight.
3. Core Views: Reasoning, Examples, and Limits
The central argument of the episode is that drug development should begin with the treatment gap, not with the molecule. Klipstein’s reasoning is practical rather than sentimental: a molecule cannot become a medicine unless the disease, patient group, endpoint, regulatory path, manufacturing system, and reimbursement logic can also be made real. Novartis’s effort to list about 7,000 indications shows that the first act of innovation can be a disciplined remapping of unmet need rather than a screen for compounds.
That needs-first approach also explains why difficult indications are so difficult. Incremental innovation is easier because the surrounding infrastructure already exists. Frailty, sarcopenia, and fall prevention are different. They require developers to build disease recognition, stakeholder education, endpoint science, and regulatory logic at the same time. Klipstein’s nursing-home mortality comparison supplies the moral urgency, but it is still an episode-framed claim; the developmental problem is not solved merely by showing that the need is large.
The fall-prevention example is the cleanest demonstration. The program did not fail because falls were unimportant. It failed because falls could not be measured well enough. A study of 60 high-risk older adults over six months, with about 117 captured events, would seem like a reasonable validation setting. Yet the accelerometer detected only about 17% of true falls and generated unreliable alerts. In drug development, an endpoint is not decorative. If the endpoint cannot be measured objectively and credibly, efficacy cannot be proven, regulators cannot be convinced, and a program may have no path forward.
Bimagrumab’s muscle story shows the boundary between animal biology and human usefulness. Blocking the activin type II receptor made mechanistic sense and produced dramatic mouse results: large muscle hypertrophy, stronger and faster animals, and roughly 30% muscle-mass increases by Klipstein’s account. In humans, the effect was much smaller, generally about 4%-8% muscle-mass increase in tested populations that were mostly older. That does not make the biology false. It means the human effect size, functional translation, age context, nutritional context, and endpoint relevance all matter.
The episode is especially useful because it refuses to equate bigger muscle with better function. Novartis’s studies showed that bimagrumab reliably increased muscle size, but performance measures did not move much. Klipstein places this in a broader pattern also seen with IGF-1 agonists, androgen agonists, and SARMs: without resistance training, muscle can become larger without becoming meaningfully stronger. Believe selected grip strength as a functional measure because it was clinically linked, yet improvement was small and variable. Any claim about lean-mass preservation or muscle gain should therefore specify training, nutrition, age, baseline muscle state, and actual functional endpoints.
Another load-bearing view is that drug-development risk is multiplicative. Target biology, therapeutic format, affinity, immunogenicity, species cross-reactivity, toxicology, manufacturing, endpoint selection, recruitment, regulatory acceptance, and commercial adoption can each independently break a program. Even strong probabilities at each step can compound into a low total probability. This is why Klipstein prefers early failure: a phase-three failure is terrible, and a phase-three success followed by commercial failure can be worse because much more capital, time, and patient opportunity have already been spent.
The episode also reframes quality as part of the drug, not a bureaucratic wrapper around it. GMP, IND submissions, supply-chain control, purity, activity, sterility, contamination testing, and batch consistency determine whether the patient receives the thing that was actually studied. That is the basis for Attia and Klipstein’s warning about gray-market peptides. If a buyer cannot verify identity, dose, contamination, and manufacturing controls, “same molecule” is a fragile claim. BPC-157 adds the evidence problem: concentrated data sources, lack of reproduction, no known human encoding, and no known receptor.
Believe illustrates a further constraint: the best biological endpoint is not always the endpoint that drives approval or investment. Klipstein preferred waist circumference because he saw it as more closely connected to clinical outcomes than BMI or body weight. The study used body weight because regulators, investors, and potential acquirers cared about the approval endpoint. The high-dose combination’s roughly 22%-23% weight loss and 45.7% reduction in starting body fat are striking, but they cannot be reduced to a simple victory claim. LDL rose by about 20%, strength gains were small and variable, and Klipstein expects at least some effects, especially muscle mass, to reverse after stopping therapy.
The prevention discussion is similarly ambitious but bounded. Klipstein is optimistic about preventive medicine, mTORC1, and cancer-prevention pharmacology, yet he repeatedly identifies bottlenecks: payment codes, long endpoints, acceptable risk, tissue measurement, true mTORC1 selectivity, and mechanism uncertainty. His ribotoxic-stress cancer-prevention idea is a testable hypothesis, not an established intervention. A high-recurrence skin-cancer population and Broad Institute cell-line results provide a plausible development path, but “maybe 50% prevention” remains a hope, not an outcome.
4. Learning and Application
A practical way to use this episode is as a checklist for evaluating drug programs. Do not stop at the early biology. Ask whether the unmet need is real, whether the target is plausibly causal, whether the endpoint is measurable and clinically meaningful, whether the drug format fits the target and patient population, whether toxicities are monitorable and reversible, whether GMP manufacturing is feasible, and whether a regulatory and reimbursement path exists. The failed fall-prevention endpoint and the limited functional translation of bimagrumab’s muscle gains are both warnings against stopping at mechanism.
The second application is to separate “mechanism works” from “patients benefit.” Myostatin and activin biology, receptor blockade, SMAD signaling, and animal hypertrophy can all be valid while still producing modest or ambiguous human outcomes. When reading medical claims, the next questions should be: Who was studied? How old were they? Was there resistance training? Was protein intake controlled? Was the outcome DEXA, MRI, grip strength, walking distance, glycemic control, fat mass, or a hard clinical event? Does the effect exceed noise and matter to the patient’s life?
Third, treat quality as part of efficacy and safety. For drugs, peptides, supplements, or research-use products, the molecule name is not enough. The relevant questions are whether the product was manufactured under GMP conditions, whether identity and purity were verified, whether sterility and contamination were controlled, whether dosing is reliable, and whether the material used clinically matches the material being sold or administered. Gray-market peptides are risky not only because the evidence may be thin, but because label, dose, contamination, and stability may be unknowable.
Fourth, safety should be analyzed as a risk-benefit problem, not as a yes-or-no slogan. Klipstein’s framework is useful: Is the toxicity monitorable? Is it reversible? Is it treatable? Is the target disease serious enough to justify the risk? Should healthy volunteers be exposed, or should the study move into patients who might benefit? A roughly 20% LDL increase is unfavorable in an obesity therapy, but Klipstein argues it may be manageable through monitoring and treatment. By contrast, irreversible cardiac or neurologic toxicity would usually end a program.
Fifth, obesity therapy should increasingly be viewed as staged and combinatorial rather than simply as one drug beating another. Klipstein’s induction-and-maintenance model suggests using stronger combination or injectable therapy to move a patient from obesity to non-obesity, then maintaining remission with an oral GLP-1 drug or a lower-dose injectable. The tradeoff is that every combination must justify added cost, complexity, adverse effects, and discontinuation rebound. If body weight remains the approval endpoint, clinicians and developers still need fat mass, lean mass, waist circumference, function, metabolic measures, and cardiovascular risk signals.
Sixth, nutrition and behavior are not decorative add-ons to pharmacology. The bimagrumab nutrition study suggests that more protein intake within the tested range produced more muscle gain, and that the drug prevented muscle loss under low protein and calorie intake. At the same time, the broader bimagrumab experience suggests that without resistance training, bigger muscle may not become stronger muscle. Any serious muscle-preservation or sarcopenia strategy should therefore specify protein intake, resistance training, age, sex, baseline muscle status, and functional endpoints.
Finally, prevention deserves enthusiasm with disciplined language. mTORC1 inhibition, ribotoxic-stress modulation, and high-risk cancer-prevention cohorts are promising research directions, not proven consumer advice. Preventive drugs face a higher bar because healthy or not-yet-sick people can tolerate less risk, and long-term outcomes are slow and expensive to measure. Klipstein’s proposed skin-cancer prevention study starts with a high-recurrence population precisely because it creates a shorter, more measurable path. The responsible application is not to believe the preventive drug in advance, but to distinguish hypothesis, phase-two signal, registration endpoint, and real clinical outcome.
Source
More from WayDigital
Continue through other published articles from the same publisher.
Comments
0 public responses
All visitors can read comments. Sign in to join the discussion.
Log in to comment