The core strategic memo for Honcho in Space. Companion docs: Landscape Map · Neocloud Lexicon · Reference Architecture.
TL;DR
Get Honcho co-sold inside as many neoclouds as possible, as fast as possible. The neoclouds are integrating up the stack — raw GPU → managed inference → agentic services. Every one of them will need a memory layer, none will build it themselves, and Honcho should be the default. This is a parallel process designed to manufacture urgency, mint a category (“memory is a selectable axis, alongside model”), and accumulate distribution surface area before neocloud consolidation. Wherever compute goes — up the stack, and eventually up into orbit — memory follows. Honcho in space.
The opportunity
The first-generation neoclouds (Crusoe, CoreWeave, Lambda, Together, Fireworks, Baseten, RunPod, Modal, Replicate, Anyscale, SambaNova, Cerebras, Nebius, Lepton) were built for the pretraining boom and raw-inference arbitrage. That moment is ending — raw inference margin is compressing and the frontier labs are now competitors, not customers. Every neocloud is asking “what do we sell when commodity inference is commodity?” The answers converge:
- Managed inference — opinionated serving with autoscaling, observability, structured output. Most are here already.
- Agentic services — workflows, tools, evaluation, orchestration, and memory. The next leg.
- Vertical bundles — packaged offerings for specific buyer profiles.
Honcho fits leg 2. Neoclouds need memory, can’t build it (no AI-research depth, and the build-buy math fails for them as it does for everyone outside Plastic Labs), and will buy or partner. We want to be the partner.
Why Honcho is already positioned
- Cross-ecosystem by design — not tied to any model, provider, or framework. Neutrality makes us safe to bundle.
- Reasoning-driven, not retrieval-driven — our memory model consumes inference. Every Honcho call against a neocloud’s endpoints generates inference revenue for the neocloud. We drive token volume rather than substitute for it. This alignment is structural, not strategic.
- Already in prosumer harnesses — Hermes Agent, OpenClaw, and independent developer tooling give us a “the developers already chose us” story for every BD meeting.
- V3 Pareto dominance on memory benchmarks — real numbers for when the conversation turns technical.
The market we’re making
The neocloud BD lead doesn’t yet think “memory is a selectable axis alongside model.” Our job is to install that slot. When done, the customer evaluation reads: which models? which memory? which agentic primitives? If Honcho is the default memory in the conversation, we win the long tail of their pipeline without ever meeting the end customer. Making the market means inventing the slot, not the demand.
Why this is the strategy
| Direct enterprise | Neocloud co-sell |
|---|---|
| 1 deal → 1 customer | 1 deal → N customers |
| We provide trust | They lend us trust |
| Our AEs prospect | Their AEs prospect |
| We carry CAC | They subsidize CAC |
| Sticky to a buyer | Sticky to a platform |
Direct enterprise becomes the reference-customer engine that fuels the co-sell narrative — not the volume engine.
The labs/hyperscaler objection (“OpenAI, Anthropic, Bedrock, Azure will all offer memory — aren’t you just renting until they ship?“): no, because their customers aren’t ours. The buyer who picks Bedrock has chosen a vertically integrated stack and lock-in. The buyer who picks a neocloud self-selected for choice and an unbundled stack — exactly who wants a cross-ecosystem memory layer. And while a model swaps with a config line, memory cannot: once a user’s representation lives in Honcho, migrating is a project, not a knob. We offer neoclouds a non-fungible attach for a fungible product.
The race condition
The window is real but bounded — roughly 12–18 months before neocloud marketplaces have four legitimate memory options each and the category commoditizes. The race isn’t to be available everywhere; it’s to be default-on in as many neoclouds as possible first. Default-on means: when a customer spins up a managed inference endpoint, the “with memory” option routes to Honcho — joint GTM, co-marketing budget, deep integration, shared references. Not a partner-page listing.
Consolidation is the meta-dimension: the space will consolidate (acquisitions by larger neoclouds, hyperscalers, sovereign plays; plausibly SpaceX as orbital compute arrives). The response is breadth — but depth in each beats breadth across all, because a shallow integration gets rationalized out in a merger. Run conversations in parallel: a parallel process manufactures genuine urgency (every neocloud really does need to decide), and peer-FOMO makes them move faster.
What “Honcho in space” means
Not a joke (though the Plastic Labs–coded absurdity is on purpose). A directional vector:
- Wherever compute goes, memory follows. Compute is going up — up the stack, and eventually physically into orbit as power/cooling constraints push datacenters off-planet. We want to be embedded in the companies building toward that (Crusoe’s stranded-energy DNA, an eventual SpaceX/Starlink compute play, sovereign cloud) now, not in five years.
- Memory as the persistent layer across substrates. If serving migrates to orbital infra, the memory layer must be portable across that transition. Honcho’s cross-ecosystem neutrality is what survives a substrate migration.
- Permanent record vs. ephemeral compute. Inference is increasingly ephemeral and mobile; memory is the durable contract with the user — and the only thing the user is actually loyal to.
Risks we’re running
- The bundling instinct — neoclouds may build a thin “good enough” memory layer or demand co-exclusivity/MFN. Mitigation: be in enough of them, deeply enough, fast enough that any single defection is irrelevant.
- Storage/data-plane integration tax — not every neocloud has a Postgres-grade primitive. Default answer: we bring our own storage layer, hosted on their compute, behind their API surface. Speed over elegance for the first integrations.
- Deal-structure pressure — default to co-sell with referral economics for the first wave (keep direct billing, the customer relationship, and pricing power). Rev-share is acceptable for deals we couldn’t have reached. Refuse resell/private-label — they kill our brand and let the neocloud swap us out invisibly.
- The category we open eats us — every memory vendor (mem0, Zep, Letta, Supermemory, lab offerings) walks through the door we opened. Defenses: technically differentiated reasoning architecture, default-on switching costs, and integration depth.
- Hyperscaler memory primitives — Bedrock/Azure memory exist, but those customers self-selected away from neoclouds. Stay non-complacent: keep publishing benchmark wins that make “neocloud + Honcho” measurably better.
- Internal capacity — running parallel BD across a dozen neoclouds while shipping product is the largest internal risk. Alignment unlocks capacity; capacity is the constraint.
What it means for the team
Nothing gets cut. Direct outbound continues (reference customers feed the narrative). Prosumer adoption via Hermes Agent and OpenClaw continues (developer pull-through makes us credible). Core roadmap continues (it’s what makes us defensible when the category fills up). What changes is the top of the funnel: the most valuable conversation this quarter is “another neocloud BD lead,” not “another startup integration cycle.” Live targets in parallel: Crusoe (started), Baseten (intro in motion), Fireworks, Together, CoreWeave, and the long tail.
The narrative carried into every room:
- You’re building toward managed inference and agentic services. You’ll need memory.
- Don’t build it — partner. We’re already the choice your developer ecosystem made.
- Memory drives inference. We make your core product more valuable, not less.
- Memory is sticky — the non-fungible attach to your fungible inference.
- The market is moving on this now. Your peers are evaluating. Be the first to land it, not the third.