Reduce. Refactor. Reuse. Loops within loops, or why reuse is not the goal.

Obvious superfluous usage of the models

First, as always, i trust everyone is safe.

Second, i have been saying something for a while that keeps coming back to me in different rooms, different codebases, different product reviews, different program reviews, and different executive conversations: Reduce. Refactor. Reuse.

That is it. Three words. No laminated framework. No twelve-box consulting bingo card. No maturity model with a gradient that looks suspiciously like it was designed by someone avoiding accountability. It sounds like an engineering mantra, which it is, but the longer i sit with it, the more convinced i am that it is not just about software. It is an operating principle for complexity. It applies to code, hardware, organizations, platforms, business processes, AI systems, manufacturing flows, partnerships, and strategy.

Most organizations get the order wrong. They start with reuse because reuse feels productive. “We already have something.” “Can we leverage the existing asset?” “Let’s standardize on the thing we built three years ago for a different customer, under different constraints, with different assumptions, and a different team that has since scattered to the wind.” That is not reuse. That is archaeology with a purchase order.

Reuse is powerful only after the thing being reused has earned the right to survive. If you reuse before you reduce, you scale clutter. If you reuse before you refactor, you scale coupling. If you reuse because a thing exists, you are not building leverage. You are distributing technical debt with better branding.

So i asked three models to react to the mantra. Not because models are authorities. They are not. Models are mirrors with a token budget. But sometimes the reflection is useful, especially when the same phrase gets interpreted through different priors. The exercise was simple: how would different thinkers react to Reduce. Refactor. Reuse. The useful part was not whether the models were “right.” The useful part was where each model placed the weight.

And yes, before somebody starts screenshotting: the Musk, Luckey, and Jobs lines below are model-generated archetypes, not real quotes. Words have meanings. So do quotation marks.

Grok: The Bar Fight Version

Grok came back with the theatrical version, imagining the mantra through Elon Musk, Palmer Luckey, and Steve Jobs. The Musk-shaped response put the weight on deletion. Not tidying. Not optimizing. Delete the part. Delete the process. Delete the line of code. Delete the meeting. Delete the requirement if the requirement is dumb. In that frame, most systems are overweight because the organization confused accumulation with progress.

“The goal isn’t elegant code. The goal is the fewest lines that still get the rocket to orbit. Everything else is cargo cult.”

That line works because it refuses to romanticize architecture. Elegant code that preserves the wrong thing is not elegance. It is embalming. Musk’s version of reduce is not “simplify the slide.” It is “prove the thing deserves to exist.” Most enterprise complexity survives because nobody wants to be the person who removes it. The system gets heavier, and then the same people wonder why it cannot move.

The Luckey-shaped response put more pressure on shipping. Reduce is the part everyone skips because writing clever abstraction feels more productive than deleting actual complexity. Refactor is where you stop lying to yourself about the design. Reuse only matters once the thing is good enough to be reused without catching fire. In hardware, this gets very real very fast. Weight, heat, power, manufacturability, maintainability, field repair, supply chain, and deployment do not care how smart the abstraction looked in the review.

“The lightest, simplest system that still works is almost always the one that ships and doesn’t catch fire.”

That is the defense-hardware version of the mantra. Reuse is not primarily about saving engineering hours. It is about reducing operational risk. A proven module, cleanly reduced and refactored, can become leverage across products, programs, and missions. But if the reusable core is dirty, then reuse becomes a force multiplier for pain.

The Jobs-shaped response treated the whole thing as product taste. Reduce until the result feels inevitable. Refactor until the cleverness disappears. Reuse only when it makes the product more human, not merely when it makes the engineer’s life easier. This is the discipline a lot of platform teams miss. The user should not be forced to admire your architecture. The experience should feel like the obvious thing that was hiding under the mess.

“The best systems disappear. If the user can feel the architecture, you failed.”

That was Grok’s useful contribution: three archetypes fighting over the same three verbs. Musk pulls toward deletion. Luckey pulls toward fielded reuse. Jobs pulls toward taste. That tension is real. It shows up in every platform conversation worth having.

ChatGPT: The Ordering Principle

ChatGPT took the more systematic route. It noticed that the mantra is deceptively simple because the order carries the philosophy. Reduce comes first because the biggest performance improvement is often deleting work, not optimizing it. Before you write code, redesign an organization, scale a process, or automate a workflow, you have to ask whether the thing should exist at all. Can the feature disappear? Can the process become unnecessary? Can we eliminate an interface, dependency, approval, or layer?

That sounds obvious right up until you sit in a meeting where everyone wants to automate a broken process instead of admitting the process is dumb. Most organizations start at the last step. They automate before they simplify. They scale before they understand. They standardize before they delete. That is how you get very sophisticated ways of preserving bad decisions.

Refactor comes next because once complexity has been reduced, what remains deserves architecture. Refactoring is not just cleaning code. At enterprise scale, it is redesigning APIs, data models, organizational seams, manufacturing flows, business processes, incentive systems, and operating rhythms. A good refactor decreases coupling while increasing adaptability. It makes the system easier to change without pretending change will stop.

Reuse comes last because most companies accidentally start there. “We already have something” becomes the opening line of a tragedy. Existing things are not automatically assets. Some are fossils. Some are scars. Some are local optimizations pretending to be platforms. Reuse only becomes powerful after reduction and refactoring. Otherwise you are simply propagating technical debt faster.

ChatGPT’s strongest line was this:

“Reuse is not the goal. Simplicity is the goal; reuse is merely the consequence of achieving simplicity.”

That is the knife. Many organizations worship reuse because it sounds efficient. But reuse without reduction and refactoring is enterprise hoarding. It creates shared libraries nobody wants to touch, common services that are common only in the sense that everyone is commonly miserable, and “platforms” that become mandatory because they are not good enough to be chosen.

The model suggested adding “repeat” to the mantra: Reduce. Refactor. Reuse. Repeat. Fair. But i pushed back. Repeat is implicit. Ordering is implicit. The strength of the phrase is that it is short enough to become instinctive. Like “ready, aim, fire,” nobody thinks you do it once and then retire to a vineyard. The repetition lives inside the discipline.

Claude: The Basis Vector Version

Claude first challenged the ordering. It argued that depending on where you stand, the honest loop might begin with reuse: what already exists, what can die, what must be cleaned, and what can then be reused by the next team. It also suggested the mantra might need a fourth R, some kind of stop condition, because otherwise refactoring can become a CTO avoiding a decision.

That annoyed me because it was partly right and partly missing the point, which is usually where useful arguments live. So i pushed back: it depends, and that is the beauty of the 3Rs.

Claude then landed on the best mathematical framing which I liked: the 3Rs are not always a fixed sequence. They are a basis. Any decision projects onto them differently depending on context. A greenfield module may weight toward reduce. A core capability used by five programs may weight toward refactor. A proven component trying to move from one program to three may weight toward reuse. Same three verbs. Different coefficients.

“A mantra is not a procedure. A procedure tells you what to do. A mantra tells you what to weigh.”

That is the part i liked. A procedure has to be correct for the situation. A mantra has to be strong enough to survive situations you did not foresee. The 3Rs do not remove judgment. They force it. They make you ask which pressure matters now: deletion, coherence, or leverage.

Claude also gave me the freediver version, which of course got my attention ( I freedive for a hobby):

“Same three strokes, but how you weight them depends on the depth, the current, and how much air you’ve got.”

That is the right metaphor. Reduce. Refactor. Reuse. Same strokes. Different water. The skill is not memorizing the order. The skill is reading the conditions without lying to yourself.

Loops Within Loops

The next step is realizing the 3Rs are not merely a sequence and not merely a basis. They are recursive. Each R contains the other two. Reduce has to be reduced, refactored, and reused. Refactor has to be reduced, refactored, and reused. Reuse has to be reduced, refactored, and reused. The loop runs inside each verb, and then the output of one loop becomes the input to the next.

That sounds like wordplay until you put it against actual work. Reducing a system is not just deletion. Good reduction has its own internal discipline. You reduce the reduce by cutting the performative requirements, zombie features, redundant approvals, decorative dashboards, and meetings that exist only because the last reorg needed artifacts. You refactor the reduce by changing the intake path so dumb requirements have fewer places to hide next time. You reuse the reduce by turning the deletion pattern into a reusable operating habit: better design reviews, better product gates, better pre-mortems, better engineering judgment, and better permission to say “no” before the system gets fat again.

Refactor has the same internal loop. You reduce the refactor by refusing to clean everything just because it exists. Some code should not be refactored. Some processes should not be redesigned. Some tools should not be modernized. Some organizations should not be optimized. They should be removed. You refactor the refactor by improving the seams that matter: interfaces, ownership, observability, support boundaries, documentation, deployment paths, and incentives. Then you reuse the refactor when the new pattern becomes an architecture other teams can adopt without inheriting the original mess.

Reuse also contains the full loop, and this is where enterprises get into the most trouble. You reduce the reuse by asking what part of the thing is actually reusable. Not the whole system. Not the customer-specific scar tissue. Not the local naming conventions. Not the assumptions that only made sense under one contract. The reusable part might be a workflow, a schema, an interface, a model evaluation, a deployment pattern, a compliance mapping, a proof point, or a failure mode. You refactor the reuse by turning that pattern into something supportable, observable, governable, documented, and owned. Then you reuse the reuse by letting the next program start from a stronger baseline and produce telemetry that tells you whether the reusable core is actually improving.

That is the loop within the loop. Every reuse creates new evidence. That evidence should trigger the next reduction. What did the second program not need? What did the third program break? What assumption failed in the fourth customer environment? What interface kept changing? What support question appeared twice? What deployment step still required a hero? Those signals tell you where to reduce again. The loop does not end at reuse. Reuse is where the next reduce gets its evidence.

The moat is not the layer; it is the loop. Take the action, capture the outcome, attribute the authority, evaluate the result, correct the workflow, and make the next decision less uncertain than the last. The 3Rs are the same pattern applied to complexity: reduce what should not exist, refactor what must exist, reuse what has earned the right to travel, then let the evidence from reuse tell you what to reduce next.

Now Apply The 3Rs To The 3Rs

A useful mantra should survive being turned against itself. So let’s do that.  We are creating a Noumena.

Noumena (plural of noumenon) is a philosophical term meaning an object or event that exists independently of human sense perception and the mind. It describes a “thing-in-itself” rather than the thing as it is experienced, heard, seen, or felt by an observer.

Origin: The word comes from the Greek noein, meaning “to think” or “to perceive with the mind”.  Also Immanuel Kant, the philosopher popularized the concept in his Britannica guide on Noumenon as part of his work on human knowledge.

#TCTRule

Words Have Meanings.

via Dr Mathew Aldridge

Can Reduce. Refactor. Reuse. be reduced? Yes, and that is why i keep resisting the urge to add a fourth word. “Repeat” is true, but unnecessary. “Freeze” is sometimes useful, but situational. “Retire” is important, but already contained inside reduce. The three words are short enough to remember and sharp enough to cause discomfort. That is a good sign. A mantra that needs a process diagram before it can be used is not a mantra. It is another artifact looking for a meeting.

Can the mantra be refactored? Also yes. The first version reads like a sequence: reduce, then refactor, then reuse. That is still useful because most organizations start with reuse and create the mess they later call platform strategy. But Claude’s basis-vector framing improves the architecture. Sometimes the situation weights toward reduce. Sometimes toward refactor. Sometimes toward reuse. The refactor is not to abandon the order; it is to understand that the order is a default, not a prison.

Can the mantra be reused? That is the real test. If it only works for code, it is a software slogan. If it works for hardware, AI, organizations, operating models, platforms, products, and partnerships, it is closer to an enterprise heuristic. The receipt is whether people can use it in a room to make a better decision. Should we delete this requirement? Should we clean this interface? Should we productize this program artifact? Should we reuse this component or quarantine it until it stops leaking? If the 3Rs help answer those questions, they are reusable. If they become wall art, they are not.

That is the self-analysis. The mantra reduces well because it is already small. It refactors well because it can shift from sequence to basis without losing meaning. It reuses well because it travels across domains. But it also has a danger: it can become too clean. Three words can hide a lot of judgment. That is why the loop matters. The 3Rs are not a substitute for thinking. They are a forcing function for thinking.

The Failure Modes

Every R has a shadow.

Reduce can become vandalism. Some people hear “reduce” and start cutting without understanding load-bearing structure. They delete the weird exception that was actually protecting the customer. They remove the approval that existed because someone once set the building on fire. They simplify the system until it is elegant and wrong. Reduction without context is not discipline. It is austerity with a hoodie.

Refactor can become avoidance. Some people hear “refactor” and discover a bottomless cave where decisions go to die. The architecture is never clean enough. The interfaces are never stable enough. The platform is always one quarter away. Refactoring is necessary, but it can become a beautiful excuse for not shipping. A refactor that never reaches reuse is not architecture. It is therapy.

Reuse can become cargo cult. This is the enterprise favorite. A thing worked once, so now it must be a platform. A local tool becomes a standard. A program artifact becomes an offering. A prototype becomes a product. A script becomes infrastructure. A PowerPoint becomes strategy. Reuse without reduction and refactoring is how an organization scales its past mistakes while congratulating itself for efficiency.

The reason the 3Rs work together is that each one corrects the others. Reduce keeps reuse from becoming debt propagation. Refactor keeps reduce from becoming demolition. Reuse keeps refactor from becoming artisanal self-expression. The tension is the point. If one R is always winning, the system is probably lying to you.

The Enterprise Loop

At enterprise scale, the 3Rs become loops within loops because every layer of work has its own version of the cycle. A team reduces a local workflow, refactors the surviving pattern, and reuses it inside a program. The program generates telemetry, failure modes, customer feedback, compliance mappings, and deployment knowledge. A product team reduces that program-specific learning to the essential pattern, refactors it into a supported capability, and reuses it across multiple programs. A platform team then reduces the common seams, refactors the interfaces and governance, and reuses the pattern across the portfolio. The company brain captures the evidence and starts the loop again.

That is how program learning becomes product learning. That is how product learning becomes platform learning. That is how platform learning becomes operating leverage. Not by declaring reuse. Not by inventorying assets. Not by creating a portal. By running the loop until the next team starts from a stronger baseline.

This matters because large organizations love to say “reuse” when what they really mean is “please make my past decision look like a platform.” We see this in software, infrastructure, internal tools, product-led solutions, proposal content, operating models, partnerships, and AI. Someone builds a local thing under local pressure. The thing works well enough to survive the contract. Then someone declares it reusable. But it was never reduced to its essential customer value. It was never refactored into clean interfaces, support boundaries, observability, governance, documentation, and ownership. So when the next program adopts it, they inherit not a product but a fossil.

Then everyone blames adoption.

No.

The thing was not ready to be reused.

Reuse is not a label. It is a property earned through reduction and refactoring. This is especially important in a product-led services company. A program can create learning. A product can encode that learning. A platform can distribute it. But only if the learning has been reduced to the essential pattern, refactored into something supportable, and reused with enough telemetry to improve the next implementation.

Otherwise, we are not compounding.

We are copy-pasting scars.

#TCTRule

Reuse is not the goal. Reuse is the receipt.

The goal is a system simple enough to understand, clean enough to change, and valuable enough to carry forward. That is true for code. It is true for hardware. It is true for AI agents. It is true for operating models. It is true for business processes. It is true for the company brain. It is true for every platform that wants to be more than a portal with a funding line.

Reduce what should not exist. Refactor what must. Reuse what has earned it. Then listen to what reuse teaches you, because every reuse is also a test. If it works, you have evidence. If it fails, you have telemetry. If it almost works, you have the next refactor. And if nobody adopts it, you may have discovered that the thing you were trying to reuse never should have existed in the first place.

That is the loop.

Reduce the work until the truth shows. Refactor the truth until it can move. Reuse only what survives contact with another customer, program, mission, or market. Then run the loop again inside the loop you just created.

Same three strokes.

Different water.

Until Then,

𝕋𝕖𝕕 ℂ. 𝕋𝕒𝕟𝕟𝕖𝕣 𝕁𝕣. (@tctjr) / X

#iwishyouwater <- Tahiti 2026 opener. Hydro dynamic complexity at its finest.

MUZAK TO BLOG BY: “Hesitation Marks” by NIN. “Copy Of A” was amazing in concert.

Footnotes

[0] The model responses referenced here came from a prompt exercise comparing how Grok, ChatGPT, and Claude interpreted “Reduce. Refactor. Reuse.” The imagined Musk, Luckey, and Jobs responses are not quotations from those people. They are model-generated archetypes. Again: words have meanings.

[1] i kept “repeat” out on purpose. If you have to say “repeat” every time, the mantra has already become a checklist. The 3Rs are a discipline, not a one-time ceremony.

[2] “Reuse is not the goal. Reuse is the receipt.” That is probably the line i would underline twice. It is also the line most enterprise platform efforts should be forced to confront before calling themselves platforms.

[3] The “loop within the loop” framing is intentionally connected to the earlier “Headless With a Spine” argument: the real moat is not a layer but an instrumented control system that acts, captures outcome, evaluates, corrects, and improves the next decision. i love control system theory, feedback loops and most importantly complexity analysis.

[4] Yes, the 3Rs are dangerously close to the old environmental slogan. That is fine. Complexity is pollution too: it accumulates, spreads, and eventually makes the whole system harder to breathe in.

Reflections from a Parking Lot: Inattentional Blindness, Smartphones, the IG-Age Main Character Default, Garage Speeding, and the Persistent Southern Wave

Main Character Default

I used to live in a room full of mirrors; all I could see was me. I take my spirit and I crash my mirrors, now the whole world is here for me to see.”
― Jimi Hendrix

First, i trust everyone reading this is safe. The other day i was easing into a parking spot headed to the gymnasium for yet another workout, just the usual slow creep, checking mirrors, trying not to clip anyone’s mirror or the curb when somebody walked straight into the side of my car. Head down. Thumbs moving. Eyes locked on whatever glowing rectangle held their attention. They bounced off, looked up surprised, mumbled something. i rolled down the window and said, “Going Somewhere?”. No damage. No drama. Just… contact with a moving vehicle they never saw coming.

It wasn’t rage-inducing. It was clarifying. In fact, it was HILARIOUS and ABSURD! Because i’ve been noticing it everywhere lately: people stepping off curbs into crosswalks without a glance left or right, crossing busy streets while scrolling like the physical world is optional. And on the flip side—drivers in parking lots treating them like short highways, zipping through without the full left-right-left scan we were all taught in driver’s ed, or backing out blind because their attention is split between the wheel and the passenger seat phone. Low-speed zones, high inattention. It feels like we’re all moving through shared space on partial autopilot (complete autopilot possibly?).

Lo and behold, after some sleuthing ye ole noggin doctors, those Psychologists have a precise term for the phone part of this: inattentional blindness. Read that again. Wow.

It’s not new. The classic “gorilla in the room” demonstration showed how dramatically our brains filter out the unexpected when attention is narrowly focused elsewhere. But smartphones turned this from a lab curiosity into everyday infrastructure. However, monkeys are funny; this is not funny.

In 2010, psychologist Ira Hyman and colleagues ran a now-famous field study on a college campus. People walking and talking on cell phones were significantly less likely to notice a unicycling clown riding across their path—complete with wig, red nose, and unicycle—than people not on phones. Many literally walked right past the clown and, when asked afterward, had no memory of anything unusual. The phone didn’t just divide attention; it created a tunnel. Later observational work has reinforced the pattern: pedestrians actively using smartphones show measurably reduced awareness of their surroundings.

The crash data is more mixed but still sobering. Pedestrian fatalities in the U.S. have been climbing for well over a decade (a reversal of earlier safety gains), with notable elevations around the pandemic years. Multiple factors are at play—larger vehicles with higher hoods that hide more of the road ahead, infrastructure that prioritizes speed, driver speeding and impairment, and yes, distraction on both sides of the windshield. Distracted walking is estimated to contribute to a meaningful but not dominant slice of incidents. It’s rarely the only thing, but it’s a controllable amplifier in an already risky system.

What has changed since the pandemic, i suspect, is the baseline level of the distractor itself. Screen time surged during lockdowns—recreational screen use for many adults and especially kids jumped by an hour or more per day on average, with much of it persisting even after restrictions lifted. We spent years conducting large parts of life through rectangles: work, socializing, entertainment, news, comfort. The habit didn’t vanish when we re-entered physical spaces.

Notifications still ping. The feed still scrolls. The pull to “just check one thing” remains engineered to be irresistible. In transitional, high-stakes environments like streets and parking lots—where you need rapid, wide situational awareness—the cost of that tug is suddenly visible in the form of near-misses and the occasional gentle bumper-to-person collision.

And it’s not only the phones. In the Instagram age or as i call it: Duh-Gram, this era of constant self-curation, story arcs, and what people are now calling “main character syndrome,” there’s a noticeable layer of entitled, almost narcissistic maneuvering in shared spaces. i’ve watched people pull right up beside me mid-maneuver in a tight lot, eyes forward or on their own reflection in the rearview, never once making eye contact or offering the smallest wave or nod that another human being exists in the vehicle next to them. Or they’ll swoop into a spot you were clearly waiting for or angling toward, no blinker, no glance back, as if the entire parking infrastructure and everyone in it are merely scenery for their personal errand montage.

The space isn’t shared; it’s theirs by default. Public pavement becomes an extension of the personal feed rather than a place requiring basic mutual recognition and courtesy.

What really fascinates—and frankly baffles me is the speeding inside parking garages and structured lots themselves. These are environments explicitly designed for crawling speeds: posted limits are typically 5–10 mph (sometimes up to 15 mph in larger commercial structures), with tight turning radii, pillars that create constant blind spots, ramps that drop or rise abruptly, low ceilings, poor sightlines around corners, and pedestrians or other vehicles appearing from every angle without warning. Yet i regularly see people treating the long straight aisles like surface streets, accelerating down ramps as if making up lost time on a commute, or maintaining highway-like momentum through the internal intersections of the garage. Even 35 to 40 mph in those confined, echoey, pillar-studded spaces feels dangerously fast because reaction distances are measured in a handful of feet, not hundreds.

A little history note here: When myself and others were working at a place called MicroSoft (small flacid anyone?)circa 2001, we used to count people who sped in parking lots and note car and license plate. Sometimes we had a casual discussion with them. One time i was “touched” by a car that was speedind. My backpack had broken the strap and i reached down (after watching both ways) and the person careened around the corner and almost “vanished me.” We had a discussion and he no longer sped.

One in five reported vehicle crashes happens in parking lots and garages; hundreds of people die and tens of thousands are injured in them every year. Many of those incidents are low-speed on paper, but the ones involving higher speeds inside the garage multiply the damage, injury risk, and sheer chaos dramatically.

Why does this happen? i genuinely don’t fully understand it. Is it simple carryover habit from open-road driving, where momentum feels normal? Time pressure that makes even a two-minute garage traversal feel like lost productivity?

Is this how YOU feel exhilarated and alive? How boring. (of course i am talking to you – who else?)

For the most part people are not curious except about themselves.
― John Steinbeck

Overconfidence from driving the same familiar structure every day? Or does the main-character default kick in hardest here: “the garage exists to get me in and out as quickly as possible, so everyone else better clear the way or get out of my story”? Whatever the mix, it turns every other problem (phone distraction, not looking, cutting in without signals or acknowledgment) into higher-stakes events. A distracted walker or an entitled lane-swiper becomes far more dangerous when the closing speed is elevated in an environment never meant for it.

Interestingly, this basic acknowledgment hasn’t entirely disappeared everywhere. In the South, it is still common for people to wave at each other—even when driving or pulling out onto two-lane roads—and especially to offer a small nod or wave when cars are maneuvering just inches apart in tight lots or garages. That tiny gesture says “i see you” in a space where modern defaults often render the other person invisible. It keeps the interaction human rather than purely transactional or oblivious.

Actually i believe there might be ten different ways to “acknowledge the other human” in the South.

Parking lots and garages are especially weird little ecosystems. Cars move in every direction at once. Pedestrians appear between bumpers or from stairwells. Sightlines are terrible by design. Everyone is either rushing to the next thing or already mentally checked into it via their device—or into their own starring role. Speeding through them is reckless. Not looking up from the phone (or from your own main-character script) while crossing or claiming them is equally so. Both are forms of the same underlying issue: fragmented or self-centered attention in an environment that punishes fragmentation with expensive metal, injured bodies, and frayed nerves.

Meditation is a way to be narcissistic without hurting anyone.
― Nassim Nicholas Taleb

On the same day, coming out of a garage, four cars in a row where I did have the right-of-way from a road rules standpoint came in the garage, and each within inches of my car turned in front of me, with the four drivers never meeting eye contact but looking straight ahead while they were turning.

It is literally NPC (non-player character) behavior in a game.

Beyond the safety numbers, though, there’s a quieter cost. We’re practicing not being fully here and not fully registering that others are here too. Missing the small negotiations of shared space—an eye contact, a slight course correction, a muttered “sorry,” or even just the basic signal that says “i see you and i’m adjusting.” In a culture optimized for continuous partial attention and personal branding, choosing full presence plus basic courtesy in public feels almost countercultural.

The Stoics would probably call attention one of the few things truly in our control; modern platforms and the performative habits they reward make that control feel optional or even quaint. Speeding in a space literally built to force slowness reveals a deeper miscalibration: we’ve stopped reading the environment’s clear cues that say “this is not a through-road; this is a careful, shared transition zone.”

i’m not arguing we throw the phones in a lake (maybe shoot them with a full metaljacket .223 AR15 at close range…) or that every driver under 35 is a narcissist. Phones are powerful tools, maps, lifelines, and sometimes the only thread keeping us connected to distant people we love. Social media didn’t invent self-interest; it just amplifies and normalizes certain versions of it.

But maybe, in the narrow slices of life where bodies and machines share pavement, especially inside the concrete echo chambers of parking garages, we can run a quick software patch on ourselves: phone down and ego down for a moment. Look up. Signal early. Ease off the accelerator. Offer the small wave or nod when you’re two inches from another car or pulling out onto the road. Treat the parking lot or garage like the high-stakes, multi-agent, shared environment it actually is.

The automobile is the most dangerous weapon in our society—cars kill more than wars do.

~ Ray Bradbury

It won’t solve bigger design and policy problems around vehicle size, street design, or speed. Those need attention too. But it’s one of the few levers each of us can pull immediately. And it might just make the shared world feel a little less like we’re all walking (or driving) through it half-blind and half-convinced the story is only about us—and a little more like the place where people still occasionally remember to wave.

A suggestion: look both ways—and up. Offer The Wave. Acknowledge the extras in your scene. And for the love of everything, ease off the gas inside the garage.

Or just keep doom scrolling when you drive. For those “Bless Yall’s Hearts, ” To no end until The End.

#iwishyouwater <- Getting the memo in Bali

Ted ℂ. Tanner Jr. (@tctjr) / X

MUZAK TO BLOG BY

Once in a Lifetime by Talking Heads.

Sing it with the following lyrics!

Because you may find yourself in a parking garage. Behind the wheel behind a large automobile, you may find yourself accelerating down a ramp you’ve taken a hundred times. You may ask yourself: why am i treating this concrete maze like an interstate when everything around me is screaming “slow down and notice the other person two inches away”? And you may tell yourself: this is not my beautiful, fully present, other-aware attention span. Wait, how did I get here? Letting the miles go by…

Note: Go watch Death Race 2000 with Charles Bronson.

Note: i saw a bumper sticker the other day – “I’m not texting i am reloading.”

Psycho-Cybernetics and The Inpert

Belief Shapes Reality

Man who says cannot be done should not interrupt Man doing.

~ December 1902 in “Puck” magazine

First, i trust everyone is safe and well into this western hemisphere summer.

Over the weekend i started ready a book with a really cool name “Psych-Cybernetics” by Dr Maxwell Maltz written in 1960. It was based on the amazing book “Cybernetics” by Dr Norbert Weiner fame.

It made me think. We’re drowning in “experts.” Every feed, every keynote, every social media post, and “powerpoint” scream specialized knowledge, certified credentials, and ten-thousand-hour mastery. And yet the breakthroughs? The real weird, world-bending stuff? It keeps coming from the edges out on the fringe.

From the people who weren’t supposed to be in the room.

Maxwell Maltz saw this in 1960. And he gave it a name that never quite went mainstream until the rest of us started needing it again: the inpert.

The Line That Stuck With Me

In the preface (page 16) to Psycho-Cybernetics, Maltz drops this quiet bomb:

Any new knowledge must usually come from the outside—not from ‘experts,’ but from what someone has defined as an ‘inpert.’

~ Dr. Maxwell Maltz

Then he lists the receipts:

  • Pasteur wasn’t an M.D. He was a chemist.
  • The Wright brothers weren’t aeronautical engineers. They built bicycles.
  • Einstein was a mathematician who torched physics.
  • Madame Curie was a physicist who changed medicine.
  • You are not your khakis.

i put the book down and started writing this blog.

Maltz wasn’t throwing shade at expertise. He was pointing at something deeper in the human success mechanism: the ability to operate outside the prescribed boundaries of a field is often where the new signal lives. Experts optimize inside the box. Inperts redraw the box from the outside.

That idea has been living rent-free in my head ever since the AI wave hit critical mass.

What an Inpert Actually Is (2026 Edition)

An inpert isn’t the opposite of competence. It’s the opposite of insularity. It’s the person who brings a mental model, a toolset, or a lived experience from one domain and slams it into another where the “experts” have stopped asking naive questions.

In actual current parlance terms, this looks like (actual humans):

  • The biologist who never took a formal machine-learning course but uses large language models and retrieval-augmented generation to compress years of wet-lab iteration into weeks—because she already knew which biological priors actually mattered.
  • The game designer who took procedural generation techniques from indie titles and applied them to climate-simulation models for coastal cities—because game engines already solved real-time physics at scale while the climate-modeling establishment was still arguing about grid resolution.
  • The lawyer who taught herself just enough agentic workflow orchestration to build intake and precedent-analysis agents that now outperform junior associates on routine work—then turned around and sold the template to mid-size firms that the big-law AI consultants hadn’t reached yet.
  • The optometrist, who had no idea how to code, started coding at the beginning of this year after a little guidance, built an electronic medical record system for his family business, and went on to build a family sovereign AI.
  • The Social Scientist who taught themselves to code and is now one of the most successful intelligence, information warfare, and PsyOps professionals in the industry.
  • A Mother and Cancer survivor who was interested in business analysis through health care data and went on to learn developer operations and deploying code at scale and now works at AWS.
  • The former government programmer who built some of the most amazing systems quit coding for a long time, got back in the saddle, and built an amazing financial retirement system.

None of these people are AI or Coding “expert” in the credentialed, conference-keynote sense. They’re inperts. They crossed the boundary carrying something the insiders had forgotten or never known.

Why This Matters More Than Ever Right Now

We’re in a moment of hyper-specialization colliding with hyper-capability. The models are getting good enough that domain depth + fresh pattern-matching beats pure model depth more often than the expert class wants to admit. The people who win are the ones willing to look stupid for a while in a new field while they port their old mental models over.

Maltz would have loved this. His whole thesis was that your self-image either expands or contracts your “success mechanism.” If you see yourself as “not an AI person,” you stay on the sidelines. If you see yourself as someone who can borrow tools and metaphors from anywhere, the mechanism lights up. The inpert identity is basically a self-image hack: I am allowed to operate outside the lines.

In fact, that is the issue. There are NO LINES. There never were any if you build new stuff or create.

The Expert Trap (and the Inpert Escape Hatch)

Experts are fantastic at refinement, safety, and scaling what already works. They’re often terrible at the first weird leap. That’s not a bug in expertise—it’s a feature of deep entrenchment. The longer you’ve defended a paradigm, the harder it is to see its edges.

Inperts don’t have that baggage. They also don’t have the credibility armor, which is why they get dismissed early and proven right later. Sound familiar? It should. It’s the same pattern Maltz documented in 1960 and we’re still watching in 2026.

How to Cultivate the Inpert Muscle

You don’t need to burn your credentials. You need to deliberately practice boundary-crossing:

  1. Pick one domain you know cold.
  2. Pick an adjacent (or wildly non-adjacent) problem you care about.
  3. Force yourself to explain the new problem using only the language and metaphors of your home domain.
  4. Build the smallest possible artifact that tests whether the translation works.

Do this in public. Document the ugly first versions. The embarrassment is the point—it’s evidence you’re operating outside the prescribed boundaries.

i’ve done versions of this my whole career—moving between big systems, startups, distributed compute, and now watching agents eat the world. Every time i felt most “not qualified,” i was usually closest to something useful that the pure experts hadn’t seen yet. Oh, the number of times i have heard – Well, that won’t work because <insert expert excuse here>.

The Ask

If you’re reading this and you’ve ever thought “I’m not technical enough” or “I don’t have the right background” or “the experts have this locked down,” I want you to hear Maltz again: new knowledge usually comes from the outside.

You are allowed to be an inpert.

The water’s warm. The models are patient. The boundaries are suggestions.

Drop me a line in the comments, tell me what you built.

Until Then,

#iwishyouwater <- footage from the June 2026 SoCal Swell. Amazing.

Ted ℂ. Tanner Jr. (@tctjr) / X

MUZAK TO BLOG BY: Peter Gabriel 3: Melt (remastered). Intruder and Games Without Frontiers surely came from the Inpert Land. The introduction to Intruder was glass on glass scratching. i applied to Real World Studios a long time ago and lost the damn rejection letter. Amazing stationery, multi colored and typed so nicely and hand-signed. Never felt so free in the face of rejection.

Book if you want to pick it up. This is the expanded version i have the original print. Psycho-Cybnernetics

Book Review: The Obstacle Is The Way

First, as always, i trust everyone is safe. Second, i haven’t written a book review in quite some time, and for the second time, i finished “The Obstacle Is The Way” , by Ryan Holiday. Thus i was compelled by the reader gods to review said tome.

It is said that before entering the sea
a river trembles with fear.
She looks back at the path she has traveled,
from the peaks of the mountains,
the long winding road crossing forests and villages.
And in front of her,
she sees an ocean so vast,
that to enter
there seems nothing more than to disappear forever.
But there is no other way.
The river can not go back.
Nobody can go back.
To go back is impossible in existence.
The river needs to take the risk
of entering the ocean
because only then will fear disappear,
because that’s where the river will know
it’s not about disappearing into the ocean,
but of becoming the ocean.


~ The River Cannot Go Back by Kahil Gibran

So, Dear Reader, you have before you this insurmountable situation, this issue, this frustration, or perhaps the looming public-speaking engagement that has already occupied far more psychological space than the event itself deserves. What if, buried somewhere inside any of the aforementioned, there is actually something beneficial for you not because every hardship arrives carrying a neatly wrapped lesson, but because almost anything can become useful once you stop allowing it to dictate the story?

i decided to re-read “The Obstacle Is The Way” by Ryan Holiday. For those who have read my blogs or actually know me, I am pro-reading a book three times. This comes from the book “How To Read A Book.’

Holiday’s central idea hit me like a wave against a sheer cliff: “What stands in the way becomes the way,” a line he pulls directly from Marcus Aurelius’ Meditations, which i had just happened to reread for the third time, this time more for comparison than consolation.

He is talking about Stoicism, of course, though not the dusty, marble-bust version taught as an elective and promptly forgotten. This is Stoicism adapted for the modern grind and organized around three disciplines: perception, action, and will. Perception is the ability to separate the event from the elaborate tragedy your ego has constructed around it; action is the deliberate movement that follows; and will is the internal architecture that remains when circumstances have removed nearly every other form of agency.

Holiday begins the discipline of perception on page 11, but the idea becomes useful around page 19 in “Recognize Your Power,” where the distinction emerges between what has happened and what we choose to make it mean. The obstacle is real; the helplessness surrounding it is often manufactured.

This becomes sharper in “Alter Your Perspective” on page 36. Perspective is not positive thinking, nor is it pretending that something objectively awful is secretly wonderful. It is the ability to change the frame without changing the facts to step far enough outside the situation that you can see its actual dimensions rather than merely experience its emotional gravity. i have started to deeply use this saying in difficult situations: “Perspective vs Perception”.

Perception and perspective are often treated as interchangeable, but they operate at different layers of the mind. Perception is the immediate interpretation of the emotional and cognitive machinery deciding what the obstacle means the moment it appears, while perspective is the distance we create afterward, the ability to step outside ourselves and see the same event within a larger system, timeline, or set of possibilities. Perception says, “This is happening to me”; perspective asks, “What is actually happening here?” One governs the first reaction, the other determines whether that reaction becomes a prison.

That is considerably more difficult than it sounds. It requires self-reflection and awareness of self.

Besides death, all failure is psychological.
—Jocko Willink

One nit for me is that the book occasionally leans too heavily into motivational-poster territory, where every failure becomes an opportunity, every rejection a redirection, and every collapsing building apparently an invitation to admire the architecture on the way down.

Not every obstacle is a conveniently positioned stepping stone. Some are institutional, irrational, or manufactured by mediocre people who protect comfortable systems; some consume years, damage confidence, and yield very little beyond the knowledge that you should never repeat the experience. Critics who dismiss the book as “airport Stoicism” are therefore not entirely wrong, because its formula can feel shallow when the problem is more serious than a missed deadline or an uncomfortable meeting.

Yet perhaps that criticism also misses the utility of the thing. A survival manual does not need to explain the complete nature of suffering; sometimes it merely needs to help you survive the night.

Rejection becomes evidence that we are unworthy. Being ignored becomes proof that we are irrelevant. An organization’s inability to use us becomes a verdict on whether we had anything useful to offer in the first place.

Rarely is any of that true. It is our ego that is talking to us, and by us, i mean the various forms of YOU.

Take, for example, ice skating. Sometimes the rink is simply too small not in its dimensions, but in what it permits you to become. Sometimes the coaches, judges, or skaters around you lack the imagination, patience, or courage to understand a different line, a different rhythm, or a program that does not fit neatly inside the expected form. Sometimes the system is built to reward repetition, protect hierarchy, and favor the technically familiar over the artistically dangerous. And sometimes you step onto the ice believing you were invited to skate your own program, only to discover that everyone else expects you to trace the same circles, hit the approved marks, and never venture too far from the boards.

Personally i prefer artistically dangerous.

Holiday’s “Live in the Present Moment,” beginning on page 45, is useful here because the mind loves to convert one obstacle into an entire imaginary future. A stalled initiative becomes a ruined career; a difficult conversation becomes permanent exclusion; one emphatic NO becomes evidence that every door ahead is already locked.

The present is usually less catastrophic than the mythology we build around it.

Years ago, in “FLAW: Not Thinking Big Enough or What Is Success?”, i wrote that we almost never think big enough in our endeavors, and when we finally believe we are approaching something audacious, grandiose, or stupendous enough, the word NO begins arriving from every direction.

No, you cannot do that.

No, it has never been done.

No, the organization does not work that way.

No, the market is not ready.

No, you are not the one.

The point was not to celebrate rejection for its own sake, or to pretend that every person saying NO is secretly frightened by your brilliance. Sometimes NO is valuable information; sometimes the idea is wrong, the timing is terrible, or your execution is deficient. The discipline is learning the difference between a warning and the gravitational pull of people who have mistaken familiarity for truth.

Every NO becomes another artifact from the boundary between what already exists and what someone is attempting to create. As I wrote then: for every NO you collect, you are one step closer to your success, not because success is cosmically guaranteed, but because anything genuinely new must first survive contact with those who cannot yet see it.

However, most often you are correct in your endeavor, just making it even more audacious. When the vision remains intact after honest inspection, collect the NO. Keep moving EverForward!

Interesting note here: i used Google’s NGRAM viewer to look up the usage of the word NO. It appears to have heightened in the mid-1940s and 1980s and has declined since then. Does that mean everything is OK? Note to self, i need to research why that happened.

That is where Holiday’s “Finding the Opportunity,” on page 53, becomes more sophisticated than the usual command to remain positive. The opportunity is not always hidden inside the obstacle like a prize in a cereal box. Sometimes the opportunity is information: the NO reveals who has imagination, where the system is brittle, which assumptions are protected, what must be rebuilt, or whether the work belongs somewhere else entirely.

Both resistance and acceptance produce a signal. Knowing which to parse is the perspective.

The Stoic response is therefore not to pretend the rejection does not hurt, nor is it to remain perfectly still while someone repeatedly strikes you with a hammer. Acceptance is not resignation. Resignation says there is nothing to be done; acceptance says this is the terrain, these are the constraints, these are the NOs I have collected, and now I must decide where to place my feet.

You may not control whether someone recognizes your value, but you control whether you continue creating it. You may not control whether an organization changes, but you control whether its inertia changes you. You may not control the politics, personalities, budgets, titles, markets, or invisible forces rearranging the board while everyone insists the game is fair, but you still control the quality and direction of your next move.

God, grant me the serenity to accept the things I cannot change, the courage to change the things I can, and the wisdom to know the difference.

~ Niebuhr – edited newer version

Which brings us to page 60: “Prepare to Act.”

Perception without action is simply a more sophisticated form of paralysis. You can understand every incentive, diagnose every dysfunction, identify every character flaw, and map the entire system down to the last trembling bureaucrat, yet none of that matters unless the understanding changes what you do next.

Perhaps the obstacle forces you to sharpen the argument, exposes a weakness that confidence had concealed, or teaches you that title and authority are not remotely the same thing. Perhaps it reveals that the mountain you have spent years climbing is leaning against the wrong wall.

Or perhaps it simply teaches you when to leave.

The obstacle becomes the way because it forces a decision that comfort would have permitted you to postpone. Friction produces information; resistance creates a signal; and failure, humiliation, stalled initiative, public stumble, an executive who underestimated you, or the latest addition to your collection of NOs all contain data, assuming you can extract it without letting the noise consume you.

An event is not an identity. A mistake is not a destiny. An obstacle is not a verdict. It is the terrain.

These words, passed down from the ancients, will carry me through every adversity and maintain my life in balance.  These four words are:  This too shall pass.

I will laugh at the world.

~OG Mandino

The Obstacle Is the Way is not a deep philosophical treatment of Stoicism; it is closer to a survival guide for the modern grind, a neon sign flickering in the darkness and reminding you to keep moving not blindly, not cheerfully, and not while pretending pain is somehow beautiful, but deliberately.

The obstacle may not have appeared for your benefit; life is not always that poetic. Once it is there, however, what you make from it belongs to you.

The insult can become clarity. The rejection can become direction. The frustration can become fuel. The public-speaking engagement can reveal a voice you did not know you possessed. The institution that refuses to change can teach you never again to confuse position with power, activity with progress, or inclusion with influence.

And the roadblock may finally force the question you should have asked long ago:

Why am I trying so hard to travel in this direction?

Perhaps the obstacle is not testing your commitment to the path.

Perhaps it is telling you to choose a better one, or even a harder one, which, Oh Dear Reader, is the best way as far as i am concerned. Always choose the hardest path. The untravelled path yields what i call “firsts”. It also has the best memories and stories later on in your life.

Do The Thing and get the courage later.

~ Joseph Campell

So, Dear Reader, when you find yourself standing before the impossible thing, resist the instinct to ask why this is happening to you. Ask what it is showing you. Separate the event from the story, study the terrain, collect the NO, and take the next deliberate action while preserving enough fire within yourself to recognize the next opportunity that will manifest itself in the #EverFoward!

Until Then,

#iwishyouwater <- Humans getting the real memo. Go Live. April 2026 Massive South Swell Zicatela Beach, Mexico.

Ted ℂ. Tanner Jr. (@tctjr) / X

Muzak To Blog By: Sugar Candy Mountain, 666, groovy 1960s throwback psychedelica.

Reflections On A Swing

An Empty Swing Not AI Generated

First i trust everyone is safe. Some say we are in the coming Golden Age, the Satya Yuga, while we are emerging from the Kali Yuga, the final, darkest, and shortest of the four cosmic ages (yugas) in Hindu cosmology, often characterized by moral decline, ignorance, and spiritual disconnection. When this era ends, it will usher in a new cycle of spiritual enlightenment, known as the Satya Yuga. Some say we are in the End Of Times. Some say We are where We are.

This past month around Maytime its much like December when everyone is running around with “not enough time”. “Posts” vary from kindergarten graduations to college graduations to marriages. For those congratulations.

Ever wonder why all of a sudden we no longer have time? Maybe its because what is stealing the time is held in your hand or unfolded in front of you.

A couple of years ago, i was standing in my yard, the drenching humidity and salty breeze wrapped around me like a blanket and a deep whisper from the Charleston Harbor, watching the “skurfer swing” move back and forth in the breeze with my beloved oak tree moss. With each sway, i realized i was envisioning one of my progeny in the swing, yet they were not there, nor are they here; they are elsewhere right now doing what i call “The Thang.”

With each sway, time seemed to wipe away the evidence that i was once the center for what i call The Chaos Engine and for my two girls and my boy, their little worlds revolving around my crazy stories or antics.

The grass will always grow back. Let’em dig.

~ Sage advice for dogs or children

In recent events i was waiting for yet another plane to board. I was looking at all the people and was reminded of a blog I wrote some time ago, “Look Up And Down And All Around!

My observation is that IT has worsened. IT refers to the amount of time people spend looking at a screen. As i say, “Things usually get worse before they get worse,” and that is not a nihilist viewpoint but an objective one. My take is that everything that happens is good at some level, but the amount of work it takes to see the good varies with the purview. For instance, the new media rule is 3 seconds. In the digital landscape, consumers and swipers operate on a “three-second rule. The attention span has been reduced to a 3-second hook.

It is called the Swipe/Scroll Reflex: The brain has been rewired for instantaneous digital gratification. If a video or webpage doesn’t provide immediate entertainment, curiosity, or value, it gets swiped away without a second thought.

Heard of Artificial Intelligence? We can now seemingly create complete archetypes and digital twins of those we “love” or think we love. Given the amount of time we spend looking at glass, screens, and LEDS glowing and the ping of a dopamine hit from text, it seems to me most would now prefer the digital facade. 3-second video of whatever, whenever. 3 Second chips of a 31,536,000-second year of 2.367e+9 in 75 years on earth – YOU have plenty of swipes!

Call me old-fashioned, but i prefer what I believe to be the “Real Thing”. The analog, if you will.

The Real Thing – The touch and smell of YOUR child, YOUR significant other, YOUR lover, YOUR friends, YOUR pet, YOUR environment.

Try this: Find a silent place. Set a timer on your favorite mind-sucking device for just five minutes and turn off all the notifications (if you can…). Most are too self-absorbed or self-conscious to try this…

Find something you love the smell of, sit there in silence, eyes shut, and listen, smell, and deeply inhale. What happened?

In the dark of night
By my side
In the dark of night
By my side
I wish you were
I wish you were

~ By My Side, INXS

Which had me thinking, listening to those you truly love, those you cherish, because ultimately they save you time, not take it, are those you want around you the most. However, you and them are sitting there with their faces buried in a screen. Even IF they aren’t, it’s sitting there on the table.

How does that make YOU feel when someone can’t put their phone or screen away?

It’s the new smoking, and it’s much more dangerous, more insidious. Where hath the time gone?

When i take you in my arms gathered forever. Sometimes it feels like a dream.

~ The Space Between – Alice Bowman

Because what if one day IT is truly just a perfect video screen rendering of your loved one that acts and talks to you like the real thing, and all you have is that digital twin? Is this better? Worse? No comment because you are afraid to show your emotions?

Clearly, it’s not the same as smelling your dog’s ear or that first kiss or touch that blew your mind.

i can hear the acc/e folks (look it up) – Yes, but what if we can simulate it all, and what if we are in a simulation?

NOTE: Trust me oh dear reader your “talking” to someone who first worked in “virtual reality” in 1993 and digital audio in 1986. I read Bostrom’s book when it first came out and have spent hours studying the math. i ws there when most of this stuff was invented.

Screen Window Rain Non AI Generated

Honestly, i don’t want to know if we are.

For me, there was a time when I remember rushing everything from bedtime stories, to half looking at their drawings to “multi-tasking” with my friends, thinking the days were endless. Now, I’d give anything for one more sticky LEGO-strewn floor, a completely wrecked couch with cookie crumbs, a completely scratched up hardwood floor from dog paws or even a clogged toilet.

As of late i still make tea and but i dont go out on a dock, though now i most always forget to drink the second mason jar full of tea much akin to habits clinging like sea mist.

For now, it’s just the hum of the fridge and the echo of silence i chase with the HVAC drone or the fan’s whir, anything to drown out the quiet that’s settled in like an uninvited guest.

Satellite’s gone
Way up to Mars
Soon it will be filled
With parking cars.

~ Satelite of Love – Lou Reed

If you’re neck-deep drowning in the mess of life, raising your own, walking that new puppy, or taking care of an elderly loved one – don’t get frustrated – take a breath – don’t wish it gone. Embrace The Chaos and even increase the Entropy. Above all, tell that person, dog, or tree truly how you feel. If you have never talked to a pet or tree your missing out….

And the cat’s in the cradle and the silver spoon
Little boy blue and the man on the moon
When you comin’ home son
I don’t know when, but we’ll get together then, Dad
We’re gonna have a good time then

~ Cats In The Cradle – Harry Chapin

One day, you’ll stare at a blank wall, out a window, or an empty swing and realize that “trouble” was the map of a life fully needed, and you’d trade the world for just one more mark, stain, or yell.

Don’t paint over the memories.

Remember, the days are long and the years are short.

Until Then,

#iwishyouwater <- Nathan Florence not caring if it’s a simulation. Fast forward to 17:02 and watch. #nearlifeexperience. If i had to go it all over again, this would be it.

Ted ℂ. Tanner Jr. (@tctjr) / X

MUZAK TO BLOG by: Different songs.

Notes:

[0] The hyphens are mine. i hate that the stochastic parrots put hyphens in generated content now.

The Token Is The New Constraint

Why is it always with a hoodie and dystopian?

Measure what is measurable, and make measurable what is not so.

~ Galileo Galilei

i have sat through more versions of the same meeting in the past couple of years than i care to admit. It always opens with a slide, sometimes a graph, and nearly always the same sentence spoken with the confidence of a human (for now) who has recently discovered fire:

AI Coding is making our engineers a buh-zillion times more productive.

~ erryone errywhere

Mehbeh. Or Mehbeh Not. This is the part nobody wants to say out loud. Just like workers who act like they don’t prompt errythang to you know where and back for literally erry-thang.

We just made it ten times easier to generate noise, ship half-finished thoughts, and call the resulting churn velocity. Look, everyone, we are agile! (“Hey where did my post it drop of the ah-gill-ee board….”)

i write this from the perspective of someone who has spent the better part of (ahem many) decades pushing bits around many types of systems, and the last several years staring directly into mission-critical workloads where a bad decision does not become a JIRA ticket or Github issue, it becomes a phone call at “Oh Dark-Thirty” local.

However, in the creation of mission-critical systems: Surgical. Defense. Logistics. Edge inference. Real-time orchestration across systems that do not forgive sloppy thinking. In those environments, you learn very quickly that productivity is not a V I B E. It is a measurable flow of high-quality decisions under constraint, and the minute you stop respecting the constraint, the system starts eating you.

AI-assisted coding did not repeal that law. It rewrote the constraint while you were grabbing another piece of pineapple ham pizza or doomscrolling.

The Unit of Work Has Changed, Quietly

Historically, the scarce resource in software was human attention. Cognitive load. Coordination overhead. The dreaded two-pizza meeting that somehow required four pizzas and resolved nothing, except everyone wondering who would take the last piece of pineapple ham. We measured productivity badly because we were measuring the wrong substrate lines of code, commits, deploys, proxies layered on top of proxies. But at least the constraint itself was stable. You had engineers. They had hours. Work flowed, or it didn’t.

Then “IT” arrived a little sooner than many of us had planned, because although WE always had hoped it would arrive, IT came in a different gift wrapping. Backpropagation was back in vogue, then came The Transformers and Deep Learning, then a pseudo-CLI where you typed a “prompt,” and then the floodgates opened: Claude Code arrived. Grok (nice model distilling there bros). Cursor (60B anyone?). A pile of agentic tooling that fundamentally changed what developers do all day. We used to spend time in deep thought, designing and thinking between compiles, but now, with the humans_still_in_the_loop, a very large fraction of the typing, scaffolding, and even architectural drafting happens inside The All Knowing Model.

Ah! Eureka! Sounds like liberation until you realize you have silently replaced one constraint with another.

This new constraint is tokens[1].

Tokens are not a fuzzy abstraction. They are a first-class engineering resource, sitting right next to our beloved CPU,GPU, memory, and network, with a dollar sign attached to each. If you run a serious org, you now have a line item that looks a lot like compute spend because that is exactly what it is. And for the first time in the history of software engineering, we can draw a clean line from idea through generation through acceptance through deployed capability with a real economic cost attached to each step. That is a gift. It is also a trap because if you do not instrument it, it will quietly devour your margin and your architecture.

Tokens are evolving into a unifying primitive across the AI stack. They function simultaneously as an economic unit, where every token is billable, forecastable, and optimizable; as a scheduling unit, mapping directly to GPU time slices through prefill and decode cycles, queueing behavior, and overall throughput; and as a cognitive unit, defining the boundary of what a model can see, reason over, and retain within its context window. That, however, is only the surface layer.

Underneath, something more fundamental is taking shape: tokens are becoming the abstraction layer that unifies currency, memory, and compute. As a currency, they represent the first truly granular pricing primitive for intelligence not measured per model or per request, but per unit of reasoning, effectively per “thought fragment.” As memory, they define bounded buffers of context, where anything outside the token window is effectively forgotten unless explicitly rehydrated through retrieval or summarization. And as compute, tokens directly drive system behavior: they determine prefill workloads, which are parallel and compute-bound, as well as decode dynamics, which are sequential and constrained by memory bandwidth and latency.

tokens = currency + memory + compute abstraction

The Illusion Of Velocity

Here is what every team sees in the first ninety days after rolling out AI-assisted development or some-thang.

PR volume goes up (ah i do hope you are even tracking them please do so, Do At Least Some-Thang). Cycle time appears to drop. Engineers report feeling faster and to be fair, they are, in the same way that a cyclist going downhill is faster than one going uphill. Leadership sees the dashboard, nods sagely[2], and declares the transformation a success. Someone updates the dreaded disease: the slideware.

Then you look one layer down, and the picture changes.

Rework climbs acting a whole lot like refactoring to somewhere? Review latency balloons because humans are now the bottleneck in a pipeline that used to be bottlenecked by writing. Architectural drift accumulates in places nobody is watching, because generation is cheap and correction is not. You did not accelerate delivery. You increased the rate at which unfinished thoughts enter the system. In a toy app this is fine. In a system that has to hold under adversarial load, it is a slow-motion incident waiting to page you.

This is the part that matters for anyone running mission-critical platforms: the system does not care how many tokens you burned or how many PRs you opened. It cares whether the thing worked, whether it held under stress, and whether it reduced the uncertainty of the next decision. Everything else is theater.

Productivity Is Flow Under Constraint — Still

The one model that survives every generation of tooling, from punch cards to Claude Code, is this: developer productivity is the rate at which high-quality decisions flow through a constrained system. AI does not repeal that. It compresses one segment of the pipeline generation and in doing so, it exposes every other weakness you had been quietly tolerating. The queue that used to hide behind slow typing now stands out like a sore thumb. The ambiguous ownership, masked by low throughput, now creates explicit collisions. The review process you always meant to fix becomes the single largest source of wait state in the system.

Three failure modes appear almost immediately, in the same order every time.

The first is batch-size inflation dressed up as speed. Engineers, armed with a model that will happily generate a thousand lines in a minute, begin opening larger PRs. Larger PRs review slower, hide more defects, and carry more coordination tax. Cycle time goes down for the author and up for the team. Net throughput falls, but it falls later, so nobody connects the dots.

The second is rework explosion again, recursion at its finest, a fractal symphony of if-then-again. First drafts are cheap now. Correct systems are not. When you watch the seven-day rewrite rate on AI-generated code, you often see it creep past twenty-five percent before anyone raises a hand. That is not productivity. That is paid trash. You are converting tokens into heat. Joules down the drain. Mother nature doesn’t like that, you know.

The third, and the one that tends to surprise people, is wait-state dominance. Once writing is no longer the bottleneck, every other stage of the pipeline becomes visible reviews, CI, environment provisioning, release gates, ownership ambiguity and most organizations were never designed to operate with those segments under scrutiny. A third and fourth pair of “Cross-Eyed” Eyes On Glass. The model did its job. The system around it did not.

What You Actually Measure

I have argued for years, including in CEO OKRs → CTO Metrics, that the job of the CTO is to translate business outcomes into a set of instrumented signals that behave like a control system, not a quarterly report. That argument becomes more important, not less, the moment tokens enter the stack.

There are four layers worth discussing, and they nest.

Flow is still the backbone. Cycle time from PR to production, review latency, and merge frequency the standard DORA-adjacent surface. The twist is that flow is only meaningful when normalized against token consumption. If cycle time is dropping but tokens per accepted change are rising superlinearly, you are not getting more efficient. You are subsidizing the illusion of speed with computing.

Quality is where most AI-assisted teams quietly fail. The signal i care about most is not defect count. It is rework velocity how quickly generated code gets rewritten. Anything rewritten inside seventy-two hours of landing is, by definition, an unstable artifact. If that number climbs, your model is producing plausible code that the system rejects. Catch it early, or pay for it architecturally.

Load, meaning cognitive and system friction, is the hidden layer. Wait-state ratio time a unit of work spends idle, divided by total cycle time, tells you where your pipeline is actually broken. Context-switching index tells you whether engineers are still doing deep work or have been reduced to prompt-and-approve operators. Ownership diffusion tells you whether accountability has been silently distributed into the ether.

Creativity is the one everyone waves their hands at, because it is hard, and because most measurement frameworks collapse the moment you try to quantify it. I want to take that seriously for a moment, because I think the AI era is actually the first time we have had the instrumentation to do it honestly.

Measuring Creativity Without Killing It

Creativity is not output volume. It is not tokens generated. It is not a commit count. Those are the things that look like creativity from a distance and fall apart on contact.

Creativity, as best i can define it in an engineering context, is the compression of complexity into a durable, elegant, high-impact solution. You are taking a messy problem and returning something that is smaller, clearer, more general, and more stable than what you started with. That is the thing we actually pay senior “engineers and creatives” for. It is also the thing that models cannot yet do reliably on their own and the thing that, if measured badly, we will incentivize people to stop doing entirely.

You cannot measure creativity directly, but you can measure its footprint.

Problem compression ratio asks how much scope, code, or complexity disappeared between the initial specification and the shipped solution.

Good engineers delete more than they add. Great creatives reframe the problem so that most of it never needed to be built. Please, folks, understand WHY before you design. Re-wind. Re-Read.

It is only a small step to measuring “programmer productivity” in terms of “number of lines of code produced per month”. This is a very costly measuring unit because it encourages the writing of insipid code, but today I am less interested in how foolish a unit it is from even a pure business point of view. My point today is that, if we wish to count lines of code, we should not regard them as “lines produced” but as “lines spent”: the current conventional wisdom is so foolish as to book that count on the wrong side of the ledger.

~ Dijkstra, E. (1987)

First-pass acceptance rate for both humans and model-generated changes tells you how often a proposed solution lands without substantial rework. In an AI-assisted world this is a double signal: it tells you about the author’s judgment and about the quality of the context they are giving the model.

Cross-domain contribution (aka project mobility) captures engineers who solve problems outside their usual lane. This is the single best leading indicator of durable technical leadership i have ever tracked. It does not scale infinitely, but its absence is diagnostic.

Token efficiency: tokens consumed per accepted, shipped, non-reworked change is the new one, and it is the one I find most honest. Because it ties cognition to economics in real time. If a team’s token efficiency is improving quarter over quarter, they are getting genuinely better at converting machine cognition into durable capability. If it is flat while spending is rising, you are paying for activity, not value. Busy is as Busy does they say.

The Token Budget Is CapEx Now

Treat it that way. I mean that literally.

For the first time, we can tie engineering output to a continuously metered economic cost. Not a quarterly cloud bill. Not a headcount ratio. A real-time, per-change, per-feature, per-decision cost of cognition. That is a level of instrumentation that finance organizations have been begging for since the first mainframe. We should not squander it by hiding it inside a developer tools P&L line item and never looking at it again.

I have a dream with that one pull request or that feature designed by that amazing product person, we can map it directly to the valuation of the company.

~ tctjr

The conversation at the leadership level stops being how productive is this creative, a question that was always slightly degrading and almost always wrong, and becomes how efficiently is this system converting tokens into mission-ready capability. That is a question you can actually answer, and more importantly, one you can act on without reducing human beings to a throughput figure.

OKRs That Force The Right Behavior

If you run a mission-critical engineering organization and you are serious about this, vague objectives are worse than no objectives. Being told these objectives when you are the one creating, designing, and building is even worse. You want constraints that bend behavior in a specific direction. Below is roughly what i would write for a platform team shipping into a real-world, high-consequence environment adapt to taste.

The objective is to increase the deployment velocity of mission-critical capabilities without increasing system risk or compute cost per unit of delivered value. That is the whole thing. It is not clever and it is not supposed to be. Mission-critical systems reward extreme clarity.

Underneath that, i would set key results that operate as a connected system: reduce cycle time by thirty percent, hold change failure rate below five percent, drive rework rate under fifteen percent, improve token efficiency tokens per accepted PR by twenty percent, and collapse wait-state ratio below twenty-five percent. Each of those moves a different lever, and moving any one of them in isolation will surface the tension with the others, which is exactly what you want. An OKR set that cannot be gamed by optimizing a single axis is an OKR set that is actually doing its job.

At the operational layer, the KPIs you look at weekly (even daily?), not quarterly, you want a very short list on the wall or an agent to display it on all the screens in the company: cycle time, rework rate, wait-state ratio, token cost per shipped feature, and acceptance rate of AI-generated code. Five numbers. If you cannot tell the story of your engineering organization with five numbers updated weekly, you do not have a control system; you have a reporting habit. And in a mission-critical environment, drift is not a quarterly problem. Drift compounds in hours.

The Weekly Conversation Is The Real Artifact

i want to be careful here, because the point of all this instrumentation is not to build a more beautiful dashboard. It is to change the conversation.

The right weekly conversation, with the right five numbers on the wall, sounds like this. Why did rework tick up this week is it a specific surface, a specific author, a specific model context, or something structural? Where are tokens being wasted are we paying for retries, for bad prompts, for agents stuck in loops? Which stage of the pipeline is accumulating wait and is that stage bottlenecked by people, tooling, or ownership? Are we shipping decisions, or are we generating artifacts that look like decisions?

If those questions are not being asked weekly, by someone with the authority to actually change the system, the system will drift. Quietly at first. Then all at once, usually on a weekend.

The Pattern That Keeps Showing Up

After enough cycles through enough organizations, one pattern keeps winning. The best teams are not the ones moving fastest in any one step. They are the ones where less gets stuck, less gets rewritten, and less gets wasted. That has always been true. AI did not change it. AI just made the deltas larger, faster, and more expensive in both directions.

The winners of the next five years will not be the teams that generate the most code. They will be the teams that waste the least, learn the fastest, and convert intent into reality with the highest signal-per-token they can sustain. The losers will ship more code than ever before, pay more for it than ever before, and create less value doing it than they did in 2019.

Closing Thought

We are not entering an era of AI-driven development. That framing is lazy, and it offloads the thinking to a model that is not qualified to do the thinking. What we are actually entering is an era of token-constrained, system-optimized, human-plus-machine engineering, which is a mouthful, but it is the honest description.

In that world, the constraint is no longer time. It is not even worth attention. For now, it is tokens, attention, and system friction, measured together, optimized together, and treated as a single economic object.

If you get that right, you do not just improve developer productivity. You build an organization that can continuously convert ideas into reality faster, cheaper, and with far higher confidence than whoever you are competing against. If you get it wrong, you will ship more than ever and mean less than ever.

That, at the end of the day, is the difference.

Measure flow. Kill wait states. Shrink work units. Respect the token. Everything else is just more code.

Until Then,

#iwishyouwater <- The Wedge March 2026.

Be Safe.

Ted ℂ. Tanner Jr.

Muzak To Blog By: Yamandú Costa, Vagner Cunha: Interpreta Concerto para Violão de 7 Cordas. This is a technically astounding piece of work. Amazing classical guitar. The recording is astounding.

Foonotes:

[1] If you made it down the stack of turtles this far, thank you for your time and attention. As a side note the word tokenization as it is used in the LLM parlance. The term is overloaded in several technology areas, including Token-Based Authentication (e.g., JWT): After a user logs in, the server issues an encrypted “token” (such as a JSON Web Token) that the client sends with subsequent requests. This avoids re-entering passwords. Security Tokens (Hardware): Physical devices (like USB keys, YubiKeys) that generate temporary codes (OTP) to prove a user possesses the device. Network Tokens (Payments): Card networks (VisaMastercard) replace sensitive Primary Account Numbers (PANs) with secure tokens to improve authorization rates and security. Blockchain (the word that shall not be said) and Web3: Tokens as Digital Assets  In blockchain, a token is a programmable, digital asset that lives on a pre-existing blockchain (like Ethereum), using smart contracts to define its behavior.Coinhouse Governance Tokens: Give holders voting power to dictate the future of a protocol. Utility Tokens: Provide access to a specific product or service within a platform (e.g., a token to access a decentralized storage network), and the list goes on and on. A token in a Large Language Model (LLM) is the fundamental, discrete unit of data that a model processes. Rather than reading text word-by-word, an LLM breaks text into smaller chunks (stemming, lemmatization, etc.), subwords, characters, or punctuation, which are then mapped to unique numerical identifiers. The cool kids term for this (from a long time ago) are “Embeddings”: These integer IDs are subsequently converted into vectors known as embeddings N dimensional space that captures semantic relationships. Right now, most of these models are, at best, a stochastic parrot (not to be confused with ParrotHeads from Jimmy Buffett), or as i view it, just major-league regex-ing at the core. So why call it a stochastic parrot, you ask? Thank you for prompting, Polly did want a cracker… We are transferring one language into another, and this is a very inefficient transfer function or an inefficient compression algorithm, just like computer languages. It only parrots what it is taught, with tokenization being a business model. My “hot take” is that the parrot (tokenization business model) will eventually die. However, that is another story, Mehbeh, for another time. If you really want to get into the details, math and code go here: Architecture Behind LLMs and Context Windows

[2]FWIW i hate pineapple ham pizza.

[3] i always wanted to nod sagely with a pipe, but I do not like smoking.

A Book Review: Tech Confidential – The Insiders’ Playbook for Daring Entrepreneurs

Tech Confidential Your Playbook and Insights To The Collesium

First, i trust everyone is safe. Second, I’m doing a book review for this installment.

Sometime ago, I was contacted by Dr. Denise Koessler Gosnell about being interviewed by her and her co-author, Ms. Kathryn Erickson. I said sure, why not? Wow, Oh Dear Readers. Just WOW!

Our Careers Are Fairytales but the part we are getting wrong is that we are the main characters

~ Tech Confidential

First, let’s set the context. Both of the authors hail from The South, as do i. That Southern grit? It’s in their DNA, turnin’ sweet tea into startup rocket fuel. Makes Tech Confidential feel like porch wisdom with a tech twist: polite as pie, sharp as a tack. Y’all get it-hospitality hides the hustle. The Awe Shucks Black Majik is being tossed around, but don’t let them aim it at you! Blessin’ hearts left and right! Both ladies have a deep, illustrious, successful background in technology in the coding and design trenches as well as boardrooms. NOTE: The above links to their LinkedIn profiles.

In a Universe where literally nobody knows everything, or anything with total certainty, curiosity is the only compass that really matters.

~ Tech Confidential

However, do not let the Southern Charm fool you, Tech Confidential: The Insider’s Playbook for Daring Entrepreneurs by Denise Koessler Gosnell (Dr DKG, as I call her) and Kathryn (Kat) Erickson is a no-holds-barred survival guide for the tech jungle. Let’s crack this thing open like an encrypted drive. We’ll zoom in on those “overclocked misfits” pushing tech to its extremes, how the best in the biz often chase performance highs in their off-hours, and the book’s razor-sharp psychological takedown of Silicon Valley as a chaotic Disneyland gone rogue. Spoiler: It’s equal parts exhilarating and exhausting, especially for those who have actually lived “the thing”.

Sometimes your passion can have such a stronghold that you get completely disconnected from reality.

~ Tech Confidential

The Overclocked Misfits: Tech’s Extreme Edge-Dwellers

At its core, Tech Confidential paints Silicon Valley not as some shiny startup Shangri-La, but as a “modern-day Colosseum where overclocked misfits battle for their shot at unicorn status and glory.” Ka-Boom! Right there in the intro, Denise and Kat nail it. “Overclocked” is pure genius metaphor: In hardware terms, it’s cranking your CPU beyond factory specs for that extra juice, risking meltdown for marginal gains. Translate to humans? We’re talking the wired weirdos, the brilliant oddballs who treat code as their personal Everest, scaling impossible heights while the rest of us mere mortals sip coffee (or tea) and debug Excel macros.

Do I want this job? You have to want the hour by hour pattern to life that comes with it.

~ Tech Confidential

These misfits are the ones operating at extreme levels, what we do in tech when the stakes are existential. Think: Engineers pulling 80-hour weeks to ship a product that’s “one breath away from the company’s last gasp,” as the book puts it. Or, founders navigating ego-fueled boardrooms where one wrong pivot means lights out and an end to your entire career. The extremes? It’s not just coding marathons; it’s the psychological warfare of toxic teams, where “brilliant oddballs” clash like gladiators. Denise and Kat don’t glorify it; they expose how this overclocking fries circuits: Burnout, identity crises, “Can I do this?”, and that nagging voice asking, “Who the hell have I become?”

i’ve seen it firsthand, hell, lived it. Push too hard, and you’re not innovating; you’re imploding. Not committing code? You’re going backward. Their advice? Tame the ego, embrace the misfit vibe without letting it consume you. Solid playbook stuff.

In the end, hapiness will come from how well you align your work with the rythym of your soul, not how long you endured a shi–y situation.

~ Tech Confidential

The Best Have Performance Hobbies: Fuel for the Fire

Here’s where it gets fascinating: The duo hints that the top performers these overclocked elites don’t just grind in the office; they channel that intensity into “performance hobbies” outside the Valley’s vortex. The book doesn’t list ’em out like a checklist, but it weaves in the idea that the best misfits thrive by balancing extremes with outlets that demand precision, speed, and flow. Think racing cars, shredding guitars, in a band, or crushing ultra-marathons, painting, hobbies where performance is measurable, adrenaline-fueled, and a reset from the abstract chaos of scaling startups.

Creating technology done well is a beautiful experience.

~ Tech Confidential

Why? Because tech at extreme levels is a mind game: Endless ambiguity, pivots, and failures. It’s that dopamine hit without the boardroom BS. The authors underscore this indirectly through anecdotes of leaders who survive by “Moneyball-ing the market,” analyzing plays like a sports stat geek, implying the crossover. Most of the best i’ve been asked to mentor (and yeah, Denise fits here) have these side quests: It keeps the overclock from redlining into burnout.

In fact, this is one of my interview questions: So, what do you do for hobbies? Reference Sidenote: I asked this to a yet-to-be college graduate on the eve of her bachelor’s in computer science. She answered: Commit to Apache Open Source Projects without hesitation. She was asked this question on a Thursday at 11PM EST and committed code in production (professionally) on the following Monday.

Pro tip from the book: If your hobby doesn’t recharge your batteries, it’s just another grind. Find one that performs as hard as you do and makes you dump all ambiguity and focus on the moment.

i’d rather work with a computer science intern eager to learn than someone who dismisses questions and steam rolls the conversation with their own vision. Skills can be taught. Process can be learned. But I can’t teach character.

~ Tech Confidential

Psychological Deep Dive: Silicon Valley’s Disneyland Mayhem

Now, the real meat of the book’s psychological autopsy on Silicon Valley as a twisted Disneyland. Denise and Kat don’t mince words: What looks like a magical kingdom of innovation free kombucha ( i personally hate that stuff), zip hoodies, and world-changing dreams is actually a mayhem machine (think Project Mayhem from Fight Club), a facade where the rides break down and the characters turn feral. They call it a “glossy utopia” sold to wide-eyed youth, but peel back the pixie dust, and it’s chaos: Overclocked misfits in a Colosseum of cutthroat politics, where success warps your soul faster than whatever happens out of Burning Man.

Software Engineering, the real real business of creating, is about helping humans get something done all the while juggling twenty grenades and dodging burning cars.

~Tech Confidential

Psychologically? It’s brutal. The authors dissect how this environment fosters imposter syndrome on steroids (the best horse juice money can buy), you’re always one funding round from irrelevance, transforming dreamers into hardened cynics. That “pre-Silicon Valley self” they mention? The one who might hate the post-unicorn version of you? That’s the mayhem: The illusion of endless opportunity masks the toll—toxic colleagues, relentless pressure, and a culture that “chews people up and spits them out in the name of innovation.” It’s like Disneyland after hours: The magic’s off, the animatronics are glitchy, and Mickey’s got a knife. Their analysis hits deep: Empathy as a survival tool, detoxing from ego-driven narratives, and reclaiming joy amid the wreckage. Humbled here again, i’ve opened doors, but books like this open minds to the mental minefield of the true Human Condition.

The only thing that matters is putting the fires about before the company looses another million dollars.

~Tech Confidential

Deeper Cuts, Lasting Lessons, and Impressions

In a sea of tech tomes preaching unicorn fairy tales, this one’s refreshingly human. It tackles the mental toll of impostor syndrome while creating the first of anything considered impossible, work-life imbalance, and toxic teams, and flips it into actionable plays. For daring entrepreneurs (or anyone slinging code in startups), it’s gold: Tips on networking without selling out, scaling without imploding, and leading with empathy in a world that rewards ruthlessness.

One of the most coveted roles to find in tech is being a strategic advisor. The strategic advisor has two modes: translating the lastest tech trend into money or chillin and smokin bud.

~ Tech Confidential

Pro tip: If you’re bootstrapping, leading a team, or restructuring a company in today’s Frontier Firm Culture, grab the hardcover; it’s got that tactile feel for late-night underlining, as you can see in these pictures. And yeah, the insights scale beyond the Valley; whether you’re in AI, biotech, or plain old SaaS, heck, even life, this playbook’s universal and for many a lifeline.

Tech Confidential isn’t just a read; it’s a mirror for Us overclocked souls. It reminds us that tech’s extremes forge legends but shatter spirits if unchecked. Balance with performance hobbies, question Disneyland After It Closes delusion, and thrive without the self-sacrifice.

If you’re in the trenches, dive in; it’ll save your sanity. If you are a partner of one of these misfits, read it. It will help you understand them. If you are just curious —like when they used to have bus tours of Silicon Valley, where we were in an aquarium —read it.

Your Great Idea is like their next Porsche; interesting but replaceable. (On VCs)

~Tech Confidential

However, for those who know you know – it is what We do. We cannot turn it off. As i say “Money and Code Never Sleep!” Send Us yet into another Colosseum! #EverForward!

DISCLAIMER: i make no royalties from this review or the book. To Dr DKG and Kat thank you i am so humbled. To Kat, you caught me off guard during the interview (which I’m hardly ever caught off guard) when you asked “Ted, what is YOUR exit?” i think about that question ever since you asked me. i might just pull a Jon Galt? Denise, i don’t even know what to say about being mentioned in this fantastic tome. i didn’t open the door, i just said there is a door over “yonder somewhere”. Again, so humbled, ladies. Much Gratitude to Both Of You.

Awe Shucks Yall Aint Had To Do All That.

Just Buy The Book HERE <- Click The Clicker Thing

Until then, Y’all Stay Overclocked, (but not fried).

#iwishyouwater <- first swell of the winter at Pipe.

Ted ℂ. Tanner Jr.

Muzak To Blog By: Night Visions by Ancestral Voices. Audio Shamanism at its finest.

SnakeByte[22]: MegaParse (An LLM’s Dream Parser with No Bits Left Behind)

Groks Idea of a MegaParser Fed Into An LLM

i’d rather have someone put a cigratte in my eye than parse pdfs.

~ BW (very accomplished coder) in a discussion about parsing healthtech pdfs

First, I trust everyone is safe. Second, I haven’t written a SnakeByte in a minute. If you’ve ever wrestled with a PDF that’s more fortress than file, you know, the kind where tables bleed into footnotes, images hide secrets, and your LLM chokes on the chaos, then you will appreciate this one.

Today, we’re diving into MegaParse, an open-source beast from QuivrHQ that’s built to crack open documents like a nutcracker on steroids. It’s optimized for LLM ingestion with zero information loss, turning messy PDFs, DOCXs, and PPTXs into clean, structured gold for your AI overlords. No more “sorry, Dave, I can’t parse that” moments.

If you’ve ever wired up RAG only to discover your PDF tables came out as ASCII i dont know what and your PowerPoints forgot their speaker notes, you’ve met the real villain: lossy parsing. Quivr’s Megaparse is an OSS parser that aims for no-loss conversion across PDFs, DOCX, PPTX, CSV/Excel, shipping markdown you can trust for embeddings and evals. Oh, and let’s not forget EDI specifications. No really.

Read On, Oh Dear Reader.

i stumbled on this gem while hunting for better ways to feed real-world docs into my own RAG experiments. In a world drowning in unstructured data (what is that saying about drowning in data and starving for information? Oh The Megatrends book), MegaParse isn’t just a parser; it’s a precision tool that respects the full spectrum: headers, footers, tables, TOCs, and even images. And get this: it comes in a “vision” mode that ropes in multimodal models like GPT-4o or Claude 3.5 to handle the gnarly stuff. Benchmarks show it smoking the competition with a 0.87 similarity ratio, way ahead of Unstructured's 0.59 or Llama Parser’s measly 0.33. That’s not hype; that’s math saying “this thing gets your docs.”

Why Bother? The Parser Wars Are Real

We’ve all been there: You dump a scanned report into an LLM, and out comes exploding salad. Traditional parsers mangle layouts, drop tables, or hallucinate whitespace where there shouldn’t be any. MegaParse flips the script by prioritizing fidelity no loss, period. It’s fast, free, and plays nice with LangChain, making it a drop-in for anyone building knowledge bases or chatty agents.

Key superpowers:

  • File Feast: Eats PDFs, DOCX, PPTX, TXT, Excel, CSV – you name it.
  • Content Clutch: Grabs tables, images, headers/footers without breaking a sweat.
  • Vision Boost: For the tough nuts, it calls in heavy hitters like GPT-4o to visually dissect pages.
  • Eval-Ready: Built-in benchmarking scripts to pit it against rivals. (Pro tip: Tweak evaluations/script.py and run it instant flex.)

It’s early days (still cooking table checkers, and structured outputs), but dang if it doesn’t feel like the parser we’ve been waiting for. Open source means you can fork it, fix it, or feast on it. Please be a good steward and contribute back. It is Apache 2.0 license.

Hands-On: Parsing Like a Pro

Let’s get dirty with some code. I’ll walk you through setup and a couple examples. (Assuming Python 3.11+ – because who lives in the past?)

Quick Install & Setup

Fire up your terminal:

pip install megaparse

I trust that wasn’t too difficult.

Ops notes (the stuff you’ll forget at 2am)

Containers: Repo includes Dockerfile and Dockerfile.gpu if you prefer hermetic builds.

System deps: PDFs/images benefit from Poppler and Tesseract; macOS also needs libmagic. Homebrew: brew install poppler tesseract libmagic.

Keys: Vision path needs an LLM key (OpenAI/Anthropic). Plain parser path can run without, depending on your inputs and slap it in a .env file (no keys in the code, boys and girls!):

OPENAI_API_KEY=your_key_here #dont put your OpenAI key for the Anthropic Key 

Example 1: Basic Parse: Effortless Extraction

Here’s the no-frills way to crack a PDF. It spits out a structured response ready for your LLM prompt. Ok, so some of you are saying ‘What’s the big deal on PDF Shredding and Parsing?” Well, check my quote at the beginning of this blog. Historically, you had to roll your own regex and then use NLTK, for example.

from megaparse import MegaParse
import json

# Initialize the parser
parser = MegaParse()

# Parse the PDF
response = parser.load("./complex_annual_report_that_no_one_wants_to_read.pdf")

# Pretty-print the output
print(json.dumps(response, indent=2))

Output? A tidy dictionary or list with sections, text, tables all intact. Feed that to your LLM, and watch it hum.

Example2: Parse -> Chunk-> Embed

 pip install megaparse tiktoken numpy sentence-transformers
from megaparse import MegaParse
from sentence_transformers import SentenceTransformer
import tiktoken, textwrap

mp = MegaParse()
doc = mp.load("./docs/board_minutes.pdf")         # -> {"markdown", "metadata", "images"}

# naive chunking by tokens
enc = tiktoken.get_encoding("cl100k_base")
def chunks(markdown, max_tokens=400):
    buf, count = [], 0
    for para in markdown.split("\n\n"):
        tokens = len(enc.encode(para))
        if count + tokens > max_tokens and buf:
            yield "\n\n".join(buf); buf, count = [], 0
        buf.append(para); count += tokens
    if buf: yield "\n\n".join(buf)

model = SentenceTransformer("all-MiniLM-L6-v2")
texts = list(chunks(doc["markdown"]))
embs  = model.encode(texts, convert_to_numpy=True)

print(f"Ingested {len(texts)} chunks; emb shape: {embs.shape}")

In the above example, you will notice tiktoken. tiktoken is a fast open-source Byte Pair Encoding (BPE) tokenizer developed by OpenAI for use with their models. It allows you to convert text strings into tokens (numerical representations) and vice versa, which is crucial for interacting with large language models (LLMs).

Swap in your vector store of choice; the point is the markdown quality gives you cleaner chunks and better recall.

In the above example.:

<N> = how many markdown chunks your PDF becomes with the ~400-token chunker.

384 = embedding size of all-MiniLM-L6-v2.

So if your document yields 12 chunks:

Ingested 12 chunks; emb shape: (12, 384)

Example 3: Vision Mode: When Pixels Get Personal

For docs with wonky scans (anything that uses an identity) or embedded visuals (think human identification or HotDogOrNot), flip to MegaParseVision. It uses a multimodal model to “see” the page, ensuring nothing gets lost in translation.

import os
from langchain_openai import ChatOpenAI
from megaparse.parser.megaparse_vision import MegaParseVision

# Set up your vision model (GPT-4o here; swap for Claude if you're fancy)
model = ChatOpenAI(model="gpt-4o", api_key=os.getenv("OPENAI_API_KEY"))

# Fire up the vision parser
vision_parser = MegaParseVision(model=model)

# Convert with eyes wide open
response = vision_parser.convert("./scanned_presentation.pptx")
print(response)

So, how to read the performance on output:

From their README benchmark (higher is better on their similarity metric):

megaparse_vision:            0.87
unstructured_with_check:     0.77
unstructured:                0.59
llama_parser:                0.33

Use this as a starting point, always test on your corpus (contracts, clinical notes, 10-Qs). They provide a evaluations/script.py hook for plugging in your own comparisons.

NOTE: I didn’t dig into the specifics of the distance similarity functions on how that is derived; however, I am guessing it’s one of the main ones on vector output.

This bad boy achieves that 0.87 benchmark score by visually cross-checking layouts. Pro move: Chain it with LangChain for RAG: parse once, query forever.

The output of the vision parser code using MegaParseVision depends on the input file (in this case, scanned_presentation.pptx) and the specific content within it, as well as the multimodal model used (e.g., GPT-4o).

Expected Output of the Vision Parser Code

The MegaParseVision class in the provided code processes the input file (a PowerPoint presentation, .pptx) using a multimodal model to extract content with high fidelity, including text, tables, images, and layout details. The output is typically a structured Python object (likely a dictionary or list) containing the parsed content, optimized for LLM ingestion. Here’s a breakdown of what you’d generally get:

Structured JSON-like Output: The response from vision_parser.convert(“./scanned_presentation.pptx”) is a structured data format (e.g., a dictionary) with keys representing different elements of the document, such as:

  • Text: Extracted text from slides, headers, footers, or annotations.
  • Tables: Structured data from any tables, often as lists or dictionaries representing rows and columns.
  • Images: Either embedded image data (e.g., base64-encoded) or references to extracted images, depending on configuration.
  • Metadata: Details like slide numbers, page layout, or document properties.
  • Visual Elements: For scanned or image-heavy documents, the vision model (e.g., GPT-4o) interprets visual content, so you might get descriptions of charts, diagrams, or other non-text elements.

Example Output Structure

Here’s a hypothetical example of what the output might look like for a simple PowerPoint slide deck with text, a table, and an image:

{
  "document_type": "pptx",
  "slides": [
    {
      "slide_number": 1,
      "text": "Welcome to Our Presentation\nKey Points:\n- Project Overview\n- Goals",
      "images": [
        {
          "description": "Company logo in top-right corner",
          "base64": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUg..."
        }
      ],
      "tables": []
    },
    {
      "slide_number": 2,
      "text": "Sales Data Q1 2025",
      "images": [],
      "tables": [
        {
          "rows": [
            ["Region", "Sales", "Growth"],
            ["North", "500K", "5%"],
            ["South", "300K", "3%"]
          ]
        }
      ]
    }
  ],
  "metadata": {
    "total_slides": 2,
    "file_name": "scanned_presentation.pptx",
    "parsed_with": "gpt-4o"
  }
}

Key Characteristics of the Output

  • Comprehensive: Includes all extractable elements (text, tables, images, etc.), leveraging the vision model to interpret scanned or visually complex content.
  • Structured for LLMs: The output is clean and organized, making it easy to feed into a language model or a RAG pipeline via LangChain.
  • Vision-Enhanced: Since MegaParseVision uses a multimodal model, it can describe images or interpret layouts that standard text parsers might miss (e.g., text embedded in images or non-standard table formats).
  • File-Specific: The exact content depends on the .pptx file’s structure. A scanned document might lean more on image descriptions, while a native PPTX might have cleaner text and table data.

Why the Output Varies

The output hinges on:

  • File Content: A text-heavy PPTX will yield more text fields; a scanned PDF converted to PPTX might emphasize image descriptions.
  • Model Choice: GPT-4o might prioritize different details compared to Claude 3.5, affecting how visual elements are described.
  • Configuration: If you’ve tweaked MegaParseVision settings (e.g., via custom prompts or parameters), the output format might differ slightly.

Bonus: API Mode for the Lazy Devs

Hate scripting? Spin up a local server at localhost:8000. Hit the /docs endpoint for Swagger-style bliss coding. Upload files, get parses zero boilerplate.

Wrapping the Byte: Parse Smarter, Not Harder

MegaParse is a reminder that good tools don’t just work; they respect your data. In the LLM era, where garbage in means garbage out, this is your anti-garbage shield. Star it, fork it, build on it and if you’re tweaking those evals, drop me a line on what you find.

NOTE: Benchmarks via their eval script run your own to confirm. No affiliation, just a fan of clean code.

NOTE: On the github “star growth” plot, they use the XKCD Python plotting library. i actually did a SnakeByte on that years ago, love the humor.

That’s your SnakeByte for today. Happy parsing!

Until Then,

Stay curious.

#iwishyouwater <- Opening Day at The Pipe 2025

Ted ℂ. Tanner Jr. (@tctjr) / X

Muzak To Blog By: Wisdom Of Clowns by Doctors Of Space. Synth meets Doom Metal meets Ambient.

CEO_OKRs_2_CTO_Metrics

Table Courtesy Of Eric Partaker


Certainty = Clear goal × Defined timeframe × Focused execution

~ SMART Goal Framework

Hello, Oh Dear Readers! First, i hope everyone is safe. Second, I’m back at the keyboard and sippin’ on that Carolina Cold Brew (iced tea) while the Palmettos sway like they’re jammin’ to some Allman Brothers or Grateful Dead (aka noodle dance).

I came across the above image in a blog by Eric Partaker. Here is the link from the originating source: OKRs For CEOs .

Eric Partaker lays out 18 CEO KPIs (Key Performance Indicators) to track as a successful company. As a developer-first CTO who’s attempted to wrangle the multi-headed hydra technical beasts at various huge as well as nascent startups, I’ve always seen tech as the engine room powering the whole ship. So, I’ve remapped these CEO metrics to CTO turf, zeroing in on how we track engineering velocity, system resilience, and innovation to drive those business outcomes.

The framework, if you will, is this:

CEO_Sets_North_Star_Company->CEO_OKRs->CEO_OKRs->CTO_KPIs->CTO_Metrics

OKRs (Objectives and Key Results) are a goal-setting framework focused on achieving ambitious, directional goals, while KPIs (Key Performance Indicators) are specific metrics used to track progress and performance. Essentially, OKRs provide the “what” and “how” of achieving a desired outcome, while KPIs provide the “how much” to measure progress.  This is also affected by the type of company, for instance, it is extremely difficult if a company has say >90% tied to strict service contracts and staff augmentation to drive this type of behavior, as you are at the mercy of the deliverable and usually a capitated margin.

But here’s the real meat: tracking ain’t just about slapping numbers on a dashboard, it’s about disaggregating the chaos, like splitting LLMs across GPUs for low-latency wins. In a general sense, “disaggregating the chaos” (made-up term) refers to the process of taking a seemingly disordered or unpredictable situation, system, or dataset and breaking it down into its smaller, individual components or elements in order to understand its underlying structure and identify patterns or causes.  NOTE: These are more behavioral mappings and metrics and are an adjunct to your real performance of your systems, although I do mention uptime and the like within this context mapping. These will be adjunctive to your core engineering and coder metrics.

We’ll dive deep into tooling like Mixpanel for user-centric product flows (think behavioral analytics on steroids), prometheus for scraping those raw infrastructure metrics (exposing endpoints for time-series data), and grafana for visualizing it all in real-time dashboards that scream actionable insights. Add in all of your engineering metrics and you have “O11y” heaven! Or for some, the other place, because Oh Dear Reader, logging all the metrics leaves no stone unturned. Also, for those that perform R&D Capitalization (if you don’t, you should), this makes the entire process brain-dead even more so than it actually is in most cases.

i’ll weave in how we’d instrument each CTO metric across these prometheus for the low-level scrapes, mixpanel for event-driven user journeys, and grafana to glue it with alerts, panels, and SLO  Service Level Objectives queries. Imagine querying prometheus for uptime histograms, funneling mixpanel events for adoption funnels, then grafana-ing it into a unified view with annotations for incidents.

We’ll assume a Kubernetes-orchestrated setup here, ’cause scale’s everything, right? Let’s break it down, OKR,KPI and Metric, with that deeper tracking lens. NOTE: If you want a Cliff’s Notes version, i made a lovely short table. Doom Scroll Oh Dear Reader, to the end.

  1. Revenue Growth Rate → Time to Market / Development Cycle Time
    Look, faster launches mean capturing market waves before they crash—I’ve seen AI models go from lab to live in weeks, spiking revenue like a Black Sabbath riff. Track this with prometheus scraping CI/CD pipeline metrics (e.g., expose /metrics endpoints for build durations, deployment frequencies via kube-state-metrics), mixpanel logging feature release events tied to user cohorts (e.g., track ‘feature_deployed’ events with properties like cycle_time), and grafana dashboards plotting histograms of lead times with alerts if cycles exceed SLOs (query: histogram_quantile(0.95, sum(rate(cycle_time_seconds_bucket[5m])) by (le))). This setup lets you correlate dev velocity to revenue spikes, spotting bottlenecks in real-time.
  2. Gross Margin → Cloud Resource Utilization
    Overprovisioning clouds is like burning cash on a bonfire—optimize it, and margins soar. We measure utilization as (allocated resources / total capacity) * 100. Prometheus shines here, scraping node-exporter for CPU/memory usage (e.g., rate(container_cpu_usage_seconds_total[5m]) / machine_cpu_cores), while mixpanel could tag resource spikes to user actions (e.g., event ‘resource_spike’ on high-traffic features). Grafana visualizes it with heatmaps of utilization over time, overlaid with cost annotations from cloud APIsset up queries like avg_over_time(node_memory_MemAvailable_bytes[1h]) to flag waste, tying back to margin erosion.
  3. Net Profit Margin → Cost Per Defect
    Defects are silent profit killers; track ’em as total fix costs / defect count. Prometheus scrapes app-level metrics like error rates (e.g., sum(rate(errors_total[5m]))), mixpanel captures user-reported bugs via events (e.g., ‘defect_encountered’ with severity props), and grafana panels trend cost-per-defect with log-scale graphs (query: sum(defect_fix_cost) / count(defects_total)). i’ve used this in in past lives to slash rework by 40%, directly padding profits add SLO alerts for defect density thresholds.
  4. Operating Cash Flow → Technical Debt Reduction
    Tech debt’s like barnacles on your hull—slows cash gen. Measure reduction as (debt items resolved / total debt) over sprints. Prometheus monitors code health via sonarqube exporters (e.g., rate(tech_debt_score[1d])), mixpanel tracks debt impact on user flows (e.g., ‘legacy_feature_used’ events), grafana dashboards with pie charts of debt categories (query: sum(tech_debt_resolved) by (type)). Chain it with burn rate queries to see cash flow correlations—personal fave: annotate debt spikes with git commit data for root causes.
  5. Cash Runway → Release Burndown
    Burndown charts predict if you’ll flame out; track as remaining tasks / velocity. Prometheus scrapes jira-like tools for burndown metrics (custom exporter for story points), mixpanel logs release milestones as events (e.g., ‘sprint_burndown_update’), grafana burndown graphs with forecast lines (query: predict_linear(release_tasks_remaining[7d], 86400 * 30)). This extends runway by flagging delays early. Especially useful in distributed systems to keep AI/ML deploys on rails without blowing budgets.
  6. Customer Acquisition Cost → Feature Usage and Adoption Rate
    High adoption turns CAC into a bargain. Measure adoption as (active users / total users) post-feature. Mixpanel owns this with funnel analysis (e.g., events like ‘feature_viewed’ → ‘feature_engaged’), prometheus for backend load from adopters (rate(feature_requests_total[5m])), grafana cohorts panels (query: sum(mixpanel_adoption_rate) over_time[30d]). Tie it to CAC by overlaying acquisition channels—deep dive: use grafana’s prometheus mixin for alerting on adoption drops below 20%. Of course, one must have an initial CAC even to log this process. Many companies have an idea of how much CAC is for a given customer or even at all. This is an imortant number for top of the funnel enterprise value chain.
  7. Customer Lifetime Value → Uptime/Downtime Rate
    Uptime’s the glue for LTV—downtime kills loyalty. Track as (total time – downtime) / total time. Prometheus is king for scraping blackbox exporters (up{job="service"}), mixpanel events for user-impacted outages (e.g., ‘downtime_experienced’), grafana SLO burn rate dashboards (query: 1 - (sum(up[1m]) / count(up[1m]))). I’ve seen this boost LTV by 25% in healthcare APIs add heatmaps for downtime patterns correlated to churn events. In past lives i posted our up time every week twitter and linkedin. “six nines” in some cases. Customers loved it.
  8. LTV-to-CAC Ratio → Automated Test Coverage
    Coverage ensures quality without tanking LTV. Measure as (tested lines / total lines) * 100. Prometheus scrapes coverage tools like istanbul where: (rate(test_coverage_ratio[1d])), mixpanel for post-deploy stability events, grafana line graphs with thresholds (query: avg(test_coverage)). Balance ratio by alerting on coverage dips—pro tip: integrate with prometheus’ recording rules for LTV projections based on quality metrics.
  9. Net Revenue Retention → System Scalability Index
    Scalability prevents revenue leaks. Index as (peak load handled / baseline) with stress tests. Prometheus scales via node_load1 (helps you understand the overall workload on a node, indicating potential resource pressure) and horizontal_pod_autoscaler, mixpanel for user growth events, grafana capacity planning panels (query: sum(rate(requests_total[5m])) / max(capacity)). This preserves NRR by forecasting breaks used it in Watson to handle surges without churn.
  10. Churn Rate → MTTR (Mean Time to Recover)
    Quick MTTR curbs churn. Calculate as sum(recovery times) / incidents. Prometheus alerts on incident durations (histogram_quantile(0.5, rate(mttr_seconds_bucket[5m]))), mixpanel ‘recovery_noticed’ events, grafana incident timelines with annotations. Deep: Set up grafana’s prometheus datasource for MTTR trends tied to churn cohorts slashed churn 15% in past gigs.
  11. Avg. Revenue Per Account → Innovation Pipeline Strength
    Pipeline fuels ARPA via upsells. Strength as (ideas in pipeline / velocity). Mixpanel tracks idea-to-feature funnels, prometheus for R&D resource metrics, grafana kanban-style boards (query: count(innovation_items) by (stage)). Visualize pipeline health to predict ARPA lifts love the fractal-like patterns in innovation flows. You can predict in some cases three months out.
  12. Burn Multiple → Code Deployment Frequency
    Frequent deploys tame burn. Frequency as deploys/day. Prometheus scrapes gitops metrics (rate(deploys_total[1d])), mixpanel for deploy-impact events, grafana frequency histograms. Correlate to burn: query sum(burn_rate) / avg(deploy_freq) keeps multiples low while accelerating ARR.
  13. Sales Cycle Length → Average Response Time
    Snappy responses shorten cycles. ART as p95 latency. Prometheus http_request_duration_seconds, mixpanel ‘response_delayed’ events, grafana latency heatmaps (query: histogram_quantile(0.95, rate(http_duration_bucket[5m]))). Tie to sales funnels for cycle reductions—game-changer in demos.
  14. Employee Turnover Rate → Team Attrition Rate
    Direct mirror; track as (exits / headcount) quarterly. Mixpanel for engagement surveys (events like 'team_feedback'), prometheus for workload metrics (e.g., oncall_burden), grafana attrition trends with forecasts. Add cultural SLOs high attrition tanks everything, as i’ve learned the hard way.
  15. Net Promoter Score → Customer Satisfaction and Retention
    Tech usability drives NPS. Mixpanel NPS events with cohorts, prometheus for support ticket resolutions, grafana score evolutions (query: avg(nps_score[30d])). Deep cohorts: Filter by product features to predict retention.
  16. Days Sales Outstanding → Platform Compatibility Score
    Compatibility smooths collections. Score as (successful integrations / attempts). Mixpanel integration events, prometheus compatibility checks, grafana success rate panels. Reduces DSO by minimizing delays—query failure rates for alerts.
  17. Growth Efficiency Ratio → Security Incident Response Time
    Fast SIRT protects growth. Like MTTR but security-focused: sum(response times) / incidents. Prometheus security exporters (e.g., falco events), mixpanel breach-impact logs, grafana incident dashboards with SLIs. Ensures efficiency without contractions.
  18. EBITDA → Employee Turnover Rate (Tech Team Focus)
    Low tech turnover boosts earnings. Same as 14 but team-specific. Mixpanel for tech satisfaction pulses, prometheus productivity metrics, grafana turnover vs. output correlations. Impacts EBITDA via reduced knowledge loss set up queries like sum(turnover_cost) / ebitda.
  19. Revenue Per Employee (calculated as Total Revenue / Average Headcount) is the North Star. In a frontier company like the ones pushing AI boundaries, this metric isn’t just important; it’s the HOLY GRAIL AFAIC. It slices through the noise to show how efficiently your team’s crankin’ out VPH (value per headcount), spotlighting if your tech wizards are amplifying revenue or just burnin’ cycles on rabbit holes. In frontier land, where innovation’s the oxygen and scale’s the game, hit high numbers here (say, north of 500K per employee like at top AI firms), and you’re signaling hyper-efficiency, attractin’ talent and investors like moths to a flame. Low? You’re leaking potential, bogged down by silos or outdated stacks. Tracking deep-dive: Prometheus scrapes raw productivity signals such as: commits_per_engineer(rate(commits_total{team="engineering"}[1d]) / headcount_gauge), mixin’ in resource efficiency (e.g., avg(cpu_usage_per_pod) to flag idle time). Mixpanel nails the revenue tie-in with event flows (e.g., ‘feature_shipped’ → ‘user_adoption’ → ‘revenue_event’, cohorting by engineer contributions via props like engineer_id). Grafana orchestrates the symphony: Custom dashboards with EPI heatmaps (query: sum(revenue_attributable) / avg(tech_headcount[30d])), overlaid with prometheus histograms for output variance and mixpanel funnels for attribution paths. Set SLOs at 80% EPI (Error Percentage Indicator) threshold alert on dips, annotate with git blame for bottlenecks, and forecast trends with predict_linear for headcount scaling. In frontier mode, this setup’s your war room: It reveals if, for instance, that new LLM fine-tunes payin’ off per engineer-hour, ensurin’ every brain cell’s punchin’ above its weight!

Slotting this as #19 to the lineup ’cause why stop at 18 when the frontier calls for more? Keeps the engine humming!

Here it is in a lovely table for all you excel spreadheet folks:

#CEO KPICTO MetricExplanation
1Revenue Growth RateTime to Market / Development Cycle TimeFaster launches mean capturing market waves before they crash
2Gross MarginCloud Resource UtilizationOverprovisioning clouds is like burning cash on a bonfire optimize it, and margins soar.
3Net Profit MarginCost Per DefectDefects are silent profit killers; track ’em as total fix costs / defect count.
4Operating Cash FlowTechnical Debt ReductionTech debt’s like barnacles on your hull slows cash gen.
5Cash RunwayRelease BurndownBurndown charts predict if you’ll flame out; track as remaining tasks / velocity.
6Customer Acquisition CostFeature Usage and Adoption RateHigh adoption turns CAC into a bargain. Measure adoption as (active users / total users) post-feature.
7Customer Lifetime ValueUptime/Downtime RateUptime’s the glue for LTV; downtime kills loyalty. Track as (total time – downtime) / total time. I’ve seen this boost LTV by 25% in healthcare APIs
8LTV-to-CAC RatioAutomated Test CoverageCoverage ensures quality without tanking LTV. Measure as (tested lines / total lines) * 100. Pro tip: integrate with prometheus’ recording rules for LTV projections based on quality metrics.
9Net Revenue RetentionSystem Scalability IndexScalability prevents revenue leaks. Index as (peak load handled / baseline) with stress tests.
10Churn RateMTTR (Mean Time to Recover)Quick MTTR curbs churn. ProTip: Set up grafana’s prometheus datasource for MTTR trends tied to churn cohorts slashed churn 15% in past gigs.
11Avg. Revenue Per AccountInnovation Pipeline StrengthPipeline fuels ARPA via upsells. Strength as (ideas in pipeline / velocity).
12Burn MultipleCode Deployment FrequencyFrequent deploys tame burn. Frequency as deploys/day.
13Sales Cycle LengthAverage Response TimeSnappy responses shorten cycles. ART as p95 latency. Tie to sales funnels for cycle reductions game changer in demos.
14Employee Turnover RateTeam Attrition RateDirect mirror; track as (exits / headcount) quarterly. Add cultural SLOs high attrition tanks everything, as I’ve learned the hard way.
15Net Promoter ScoreCustomer Satisfaction and RetentionTech usability drives NPS. Deep cohorts: Filter by product features to predict retention.
16Days Sales OutstandingPlatform Compatibility ScoreCompatibility smooths collections.
17Growth Efficiency RatioSecurity Incident Response TimeFast SIRT protects growth. Like MTTR but security-focused: sum(response times) / incidents.
18EBITDAEmployee Turnover Rate (Tech Team Focus)Low tech turnover boosts earnings. Same as 14 but team-specific.
19Revenue Per EmployeeEngineer Productivity Index (EPI)Revenue Per Employee (calculated as Total Revenue / Average Headcount) is the North Star. In a frontier company like the ones pushin’ AI boundaries, this metric ain’t just important; it’s the holy grail. Hit high numbers here (say, north of $500K per employee like at top AI firms),

Table 1.0 Easy Explanations and Mappings

Whew, that’s the full rundown feels like paddling through a fractal wave, but with these tools, you’re not just tracking; you’re orchestrating a symphony of data!

Until Then,

#iwishyouwater <- Koa Rothman At Teachpoo Largest in 15 years. They Got The Memo.

Ted ℂ. Tanner Jr. (@tctjr) / X

Muzak To Blog By Devo “Duty Now For the Future” and “Q: Are We Not Men?” They were amazing in concert. Lyrics are so cogent for today

A Survey Of Architectures And Methodologies For Distributed LLM Dissaggregation

Grok 4’s Idea of a Distributed LLM

I was kicking back in my Charleston study this morning, drinking my usual unsweetened tea in a mason jar, the salty breeze slipping through the open window like a whisper from the Charleston Harbor, carrying that familiar tang of low tide “pluff mud” and distant rain. The sun was filtering through the shutters, casting long shadows across my desk littered with old notes on distributed systems engineering, when I dove into this survey on architectures for distributed LLM disaggregation. It’s a dive into the tech that’s pushing LLMs beyond their limits. As i read the numerous papers and assembled commonalities, it hit me how these innovations echo the battles many have fought scaling AI/ML in production, raw, efficient, and unapologetically forward. Here’s the breakdown, with the key papers linked for those ready to dig deeper.

This is essentially the last article in a trilogy. The sister survey is a blog A Survey of Technical Approaches For Distributed AI In Sensor Networks. Then, for a top-level view, i wrote SnakeByte[21]: The Token Arms Race: Architectures Behind Long-Context Foundation Models, so you’ll have all views of a complete system, sensors->distributed compute methods->context engineering.

NOTE: By the time this is published, a whole new set of papers will come out, and i wrote (and read the papers) in a week.

Overview

Distributed serving of LLMs presents significant technical challenges driven by the immense scale of contemporary models, the computational intensity of inference, the autoregressive nature of token generation, and the diverse characteristics of inference requests. Efficiently deploying LLMs across clusters of hardware accelerators (predominantly GPUs and NPUs) necessitates sophisticated system architectures, scheduling algorithms, and resource management techniques to achieve low latency, high throughput, and cost-effectiveness while adhering to Service Level Objectives (SLOs). As you read the LLM survey, think in terms of deployment architectures:

Layered AI System Architecture:

  • Sensor Layer: IoT, Cameras, Radar, LIDAR, electro-mag, FLIR etc
  • Edge/Fog Layer: Edge Gateways, Inference Accelerators, Fog Nodes
  • Cloud Layer: Central AI Model Training, Orchestration Logic, Data Lake

Each layer plays a role in collecting, processing, and managing AI workloads in a distributed system.

Distributed System Architectures and Disaggregation

Modern distributed Large Language Models serving platforms are moving beyond monolithic deployments to adopt disaggregated architectures. A common approach involves separating the computationally intensive prompt processing (prefill phase) from the memory-bound token generation (decode phase). This disaggregation addresses the bimodal latency characteristics of these phases, mitigating pipeline bubbles that arise in pipeline-parallel deployments KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving. As a reminder in LLMs, KV cache stores key and value tensors from previous tokens during inference. In transformer-based models, the attention mechanism computes key (K) and value (V) vectors for each token in the input sequence. Without caching, these would be recalculated for every new token generated, leading to redundant computations and inefficiency.

Systems like Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving propose a KVCache-centric disaggregated architecture with dedicated clusters for prefill and decoding. This separation allows for specialized resource allocation and scheduling policies tailored to each phase’s demands. Similarly, P/D-Serve: Serving Disaggregated Large Language Model at Scale focuses on serving disaggregated LLMs at scale across tens of thousands of devices, emphasizing fine-grained P/D organization and dynamic ratio adjustments to minimize inner mismatch and improve throughput and Time-to-First-Token (TTFT) SLOs. KVDirect: Distributed Disaggregated LLM Inference explores distributed disaggregated inference by optimizing inter-node KV cache transfer using tensor-centric communication and a pull-based strategy.

Further granularity in disaggregation can involve partitioning the model itself across different devices or even separating attention layers, as explored by Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache, which disaggregates attention layers to enable flexible resource scheduling and enhance memory utilization for long contexts. DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving unifies and extends both colocated and disaggregated paradigms using a micro-request abstraction, splitting requests into segments for balanced load across unified GPU instances.

The distributed nature also necessitates mechanisms for efficient checkpoint loading and live migration. ServerlessLLM: Low-Latency Serverless Inference for Large Language Models proposes a system for low-latency serverless inference that leverages near-GPU storage for fast multi-tier checkpoint loading and supports efficient live migration of LLM inference states.

Scheduling and Resource Orchestration

Effective scheduling is paramount in distributed LLM serving due to heterogeneous request patterns, varying SLOs, and the autoregressive dependency. Existing systems often suffer from head-of-line blocking and inefficient resource utilization under diverse workloads.

Preemptive scheduling, as implemented in Fast Distributed Inference Serving for Large Language Models, allows for preemption at the granularity of individual output tokens to minimize latency. FastServe employs a novel skip-join Multi-Level Feedback Queue scheduler leveraging input length information. Llumnix: Dynamic Scheduling for Large Language Model Serving introduces dynamic rescheduling across multiple model instances, akin to OS context switching, to improve load balancing, isolation, and prioritize requests with different SLOs via an efficient live migration mechanism.

Prompt scheduling with KV state sharing is a key optimization for workloads with repetitive prefixes. Preble: Efficient Distributed Prompt Scheduling for LLM Serving is a distributed platform explicitly designed for optimizing prompt sharing through a distributed scheduling system that co-optimizes KV state reuse and computation load-balancing using a hierarchical mechanism. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool integrates context caching with disaggregated inference, supported by a global scheduler that enhances cache reuse through a global prompt tree-based locality-aware policy. Locality-aware fair scheduling is further explored in Locality-aware Fair Scheduling in LLM Serving, which proposes Deficit Longest Prefix Match (DLPM) and Double Deficit LPM (D2LPM) algorithms for distributed setups to balance fairness, locality, and load-balancing.

Serving multi-SLO requests efficiently requires sophisticated queue management and scheduling. Queue management for slo-oriented large language model serving is a queue management system that handles batch and interactive requests across different models and SLOs using a Request Waiting Time (RWT) Estimator and a global scheduler for orchestration. SLOs-Serve: Optimized Serving of Multi-SLO LLMs optimizes the serving of multi-stage LLM requests with application- and stage-specific SLOs by customizing token allocations using a multi-SLO dynamic programming algorithm. SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference proposes a service-aware and latency-optimized resource sharing framework for large language model inference.

For complex workloads like agentic programs involving multiple LLM calls with dependencies, traditional request-level scheduling is suboptimal. Autellix: An Efficient Serving Engine for LLM Agents as General Programs treats programs as first-class citizens, using program-level context to inform scheduling algorithms that preempt and prioritize LLM calls based on program progress, demonstrating significant throughput improvements for agentic workloads. Parrot: Efficient Serving of LLM-based Applications with Semantic Variable focuses on end-to-end performance for LLM-based applications by introducing the Semantic Variable abstraction to expose application-level knowledge and enable data flow analysis across requests. Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution optimizes for tool-aware LLM serving by enabling tool partial execution alongside LLM decoding.

Memory Management and KV Cache Optimizations

The KV cache’s size grows linearly with sequence length and batch size, becoming a major bottleneck for GPU memory and throughput. Distributed serving exacerbates this by requiring efficient management across multiple nodes.

Effective KV cache management involves techniques like dynamic memory allocation, swapping, compression, and sharing. KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving proposes KV-cache streaming for fast, fault-tolerant serving, addressing GPU memory overprovisioning and recovery times. It utilizes microbatch swapping for efficient GPU memory management. On-Device Language Models: A Comprehensive Review presents techniques for managing persistent KV cache states including tolerance-aware compression, IO-recompute pipelined loading, and optimized chunk lifecycle management.

In distributed environments, sharing and transferring KV cache states efficiently are critical. MemServe introduces MemPool, an elastic memory pool managing distributed memory and KV caches. Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache leverages a pooled GPU memory strategy across a cluster. Prefetching KV-cache for LLM Inference with Dynamic Profiling proposes prefetching model weights and KV-cache from off-chip memory to on-chip cache during communication to mitigate memory bottlenecks and communication overhead in distributed settings. KVDirect: Distributed Disaggregated LLM Inference specifically optimizes KV cache transfer using a tensor-centric communication mechanism.

Handling Heterogeneity and Edge/Geo-Distributed Deployment

Serving LLMs cost-effectively often requires utilizing heterogeneous hardware clusters and deploying models closer to users on edge devices or across geo-distributed infrastructure.

Helix: Serving Large Language Models over Heterogeneous GPUs and Networks via Max-Flow addresses serving LLMs over heterogeneous GPUs and networks by formulating inference as a max-flow problem and using MILP for joint model placement and request scheduling optimization. LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization supports efficient serving on heterogeneous GPU clusters through adaptive model quantization and phase-aware partition. Efficient LLM Inference via Collaborative Edge Computing leverages collaborative edge computing to partition LLM models and deploy them on distributed devices, formulating device selection and partition as an optimization problem.

Deploying LLMs on edge or geo-distributed devices introduces challenges related to limited resources, unstable networks, and privacy. PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services provides a personalized inference scheduling framework with edge-cloud collaboration for diverse LLM services, optimizing scheduling and resource allocation using a UCB algorithm with constraint satisfaction. Distributed Inference and Fine-tuning of Large Language Models Over The Internet investigates inference and fine-tuning over the internet using geodistributed devices, developing fault-tolerant inference algorithms and load-balancing protocols. MoLink: Distributed and Efficient Serving Framework for Large Models is a distributed serving system designed for heterogeneous and weakly connected consumer-grade GPUs, incorporating techniques for efficient serving under limited network conditions. WiLLM: an Open Framework for LLM Services over Wireless Systems proposes deploying LLMs in core networks for wireless LLM services, introducing a “Tree-Branch-Fruit” network slicing architecture and enhanced slice orchestration.

On the of most recent papers that echo my sentiment from years ago where is i’ve said “Vertically Trained Horizontally Chained” (maybe i should trademark that …) is Small Language Models are the Future of Agentic AI where they lay out the position that specific task LLMs are sufficiently robust, inherently more suitable, and necessarily more economical for many invocations in agentic systems, and are therefore the future of agentic AI. The argumentation is grounded in the current level of capabilities exhibited by these specialized models, the common architectures of agentic systems, and the economy of LM deployment. They further argue that in situations where general-purpose conversational abilities are essential, heterogeneous agentic systems (i.e., agents invoking multiple different models chained horizontally) are the natural choice. They discuss the potential barriers for the adoption of vertically trained LLMs in agentic systems and outline a general LLM-to-specific chained model conversion algorithm.

Other Optimizations and Considerations

Quantization is a standard technique to reduce model size and computational requirements. Atom: Low-bit Quantization for Efficient and Accurate LLM Serving proposes a low-bit quantization method (4-bit weight-activation) to maximize serving throughput by leveraging low-bit operators and reducing memory consumption, achieving significant speedups over FP16 and INT8 with negligible accuracy loss.

Splitting or partitioning models can also facilitate deployment across distributed or heterogeneous resources. SplitLLM: Collaborative Inference of LLMs for Model Placement and Throughput Optimization designs a collaborative inference architecture between a server and clients to enable model placement and throughput optimization. A related concept is Split Learning for fine-tuning, where models are split across cloud, edge, and user devices Hierarchical Split Learning for Large Language Model over Wireless Network. BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models enables multi-tenant finer-grained serving by partitioning models into blocks, allowing component sharing, adaptive assembly, and per-block resource configuration.

Performance and Evaluation Metrics

Evaluating and comparing distributed LLM serving platforms requires appropriate metrics and benchmarks. The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving suggests a trade-off between serving context length (C), serving accuracy (A), and serving performance (P). Developing realistic workloads and simulation tools is crucial. BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems provides a real-world workload dataset for optimizing LLM serving systems, revealing limitations of current optimizations under realistic conditions. LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale is a HW/SW co-simulation infrastructure designed to model dynamic workload variations and heterogeneous processor behaviors efficiently. ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency focuses on a holistic system view to optimize LLM serving in an end-to-end manner, identifying and addressing bottlenecks beyond just LLM inference.

Conclusion

The landscape of distributed LLM serving platforms is rapidly evolving, driven by the need to efficiently and cost-effectively deploy increasingly large and complex models. Key areas of innovation include the adoption of disaggregated architectures, sophisticated scheduling algorithms that account for workload heterogeneity and SLOs, advanced KV cache management techniques, and strategies for leveraging diverse hardware and deployment environments. While significant progress has been made, challenges remain in achieving optimal trade-offs between performance, cost, and quality of service (QOS) across highly dynamic and heterogeneous real-world scenarios.

As the sun set and the neon glow of my screen dimmed, i wrapped this survey up, leaving me pondering the endless horizons of AI/ML scaling like waves crashing on the shore, relentless and full of promise and thinking how incredible it is to be working in these areas where what we have dreamed for decades has come to fruition?

Until Then,

#iwishyouwater

Ted ℂ. Tanner Jr. (@tctjr) / X

MUZAK TO BLOG BY: Vangelis, “L’apocalypse de animax (remastered). Vangelis is famous for “Chariots Of Fire” and “Blade Runner” Soundtracks.