Institutional-scale reasoning
Long-horizon, multidisciplinary inquiry across teams, models, and bodies of evidence. Investigate complex questions without losing provenance, contradictions, or the institution’s decision criteria.
Mainely Code’s contribution to AI science
NORTHSTAR is a research and engineering pursuit: to advance synthetic reasoning through the governed cooperation of many models and many systems, for the benefit of humankind.
The central idea
Not one model asked to know everything. Not a crowd of models repeating an answer. A coordinated lattice that can investigate, disagree, verify, and synthesize.
The vision encompasses every major general-purpose model family alongside specialist models, tools, and private institutional systems. Open and commercial models should contribute according to their actual capabilities, the mission’s policies, and the evidence they can provide—not a vendor hierarchy.
Some contributors investigate independently. Others look for flaws, seek contradictory evidence, test calculations, or compare hypotheses. The executive coordinates that work; people retain authority over purpose and consequential action.
The desired result is a level of synthetic reasoning beyond what any isolated participant can provide. That is a research objective to evaluate, not an achievement asserted by this website.
The intended scale begins with dozens or hundreds of participating nodes. Thousands are a longer-term horizon, contingent on advances in control-plane hardware and demonstrated coordination capacity—not a claim about today’s deployment size.
Three dimensions of the vision
Long-horizon, multidisciplinary inquiry across teams, models, and bodies of evidence. Investigate complex questions without losing provenance, contradictions, or the institution’s decision criteria.
Coordinate everything from a stack of old laptops to private clusters and queued GPU time across major public clouds. Grow, shrink, or redistribute work as authorized capacity changes.
Retain deliberate control of data, model access, infrastructure, identity, keys, and policy. A private or sovereign mission must not depend on quietly sending protected work somewhere else.
A scientific standard
Coordination has to earn its complexity. The research program should compare a lattice-assisted workflow with strong single-model and simpler multi-model baselines under comparable constraints.
Useful evaluation asks whether the system improves evidence coverage, discovers meaningful contradictions, calibrates uncertainty, resists correlated mistakes, and produces better-supported conclusions. Latency, compute, energy, and cost belong in that assessment too.
Repeated agreement is not independence. More tokens are not proof. A visually impressive lattice is not a reasoning benchmark.
The evidence we intend to requireFor the benefit of humankind
We want to further the science of AI in ways that help people and institutions understand difficult problems more deeply, examine consequences more carefully, and make more responsible decisions.
Scientific discovery. Resilient food and resource systems. Energy and climate research. Complex public-interest questions. These are horizons for inquiry—not promises that software alone can resolve humanity’s challenges.