One can be anywhere. The other has to be near you.
Follow one answer from the place the model is built to the moment it reaches you, and see why the last stop has to be close.
Start at step oneA language model is software that has learned the patterns of language well enough to answer in it. It is built once, then used again and again. The first two steps happen far away, long before you ask anything. The last two happen close by, in less time than a blink.
Where: a data center, wherever land and power are plentifulWhen: months before you ask
The model starts out knowing nothing. In a data center it is given an enormous body of text and one exercise, repeated billions of times: guess what comes next, check the guess, adjust. Thousands of processors run that exercise together for weeks or months. The process is called training.
GatherA vast library of text is assembled: books, articles, websites, code.
TrainThe model practices on it until it has learned how language works.
FinishWhat comes out is the model: a file that holds everything it learned.
Where: over long-distance fiberWhen: once for each model
A finished model is a file, and a file can be copied. Copies travel to Inference Points across the country, each one placed inside the area it will serve. It is the only long trip the model makes. From then on it stays put, loaded and ready, a short distance from the people who will use it.
Where: an Inference Point, inside your communityWhen: the moment someone asks
Now someone asks. A person types a question, a scanner sends an image, a machine on a production line sends a reading. The request reaches the Inference Point over local fiber, the model computes an answer on the spot, and the answer goes back. The process is called inference.
AskA request arrives from a person, a business or a machine nearby.
ComputeThe model works out a fresh answer. Nothing is pulled from storage.
DeliverThe answer goes back to whoever asked, milliseconds later.
Where: back with youWhen: milliseconds later
A signal in fiber covers roughly 125 miles each millisecond, and no engineering makes glass faster. Every mile is paid for twice: once out, once back. With the model nearby, the trip is too short to notice. With the model on a distant campus, the trip is most of the wait.
Where is the model running?
1 to 5 msfrom question to answer
30 to 80 msfrom question to answer
Both bars share one scale, measured end to end.
Work that can wait is fine at a distance. Summarizing a report overnight does not care about the round trip. Work that happens in real time does.
A conversationA voice agent that answers before the pause turns awkward.
A production lineA camera that rejects the bad part before it reaches the next station.
An emergencyA dispatch map that moves when the ambulance does.
Building a model and using one are different work, done in different places, on different clocks. So the infrastructure behind each is built, placed and measured differently.
| Building the model, in a data center | Using the model, at an Inference Point | |
|---|---|---|
| It is called | Training | Inference |
| How often | Once for each model | Every time anyone asks |
| How long | Weeks to months | Milliseconds |
| Who is waiting | No one | Someone, every time |
| Where it has to be | Anywhere land and power are plentiful | Near the people it answers |
A data center exists to house and store. An Inference Point exists to compute and deliver, locally.
A difference of purpose, not size.
No. A small warehouse is still a warehouse. What separates the two is what each is for and where it has to be. Move a data center 500 miles and its customers never notice. Move an Inference Point 500 miles and it stops doing its job.
Data centers do essential work. Without them there would be no model to run, and nowhere to keep what the world stores. An Inference Point does the other half: it puts the finished model within reach of the people, businesses and machines that need an answer now.
The systems a community cannot do without share one trait: each is built within reach of the people it serves. As AI moves from something people consult to something that acts in real time, the compute behind it joins that list. The Edge calls it mission-critical infrastructure, because the work that depends on it cannot pause.
See who uses one, and howSubstationSteps power down for the streets around it.
Cell towerStands where its signal can reach you.
Fire stationSits where its crews can reach you in time.
Inference PointComputes where its answers can reach you in time.
An Inference Point is installed inside an existing commercial building, on the distribution lines that already serve the street. No campus, no cleared land, no new construction.
See how, who uses one, and the growth a community can expect.
See the community benefitsThe Edge is building a nationwide network of Inference Points that connects AI compute with people, businesses, and machines.