Atm the meta for local inference is unified memory and “routed” local agents (multiple smaller role-specific agents in a trench coat)
The former is standout for cost efficiency (e.g., 4x RDMA 48gb Minis for a 192gb cluster @ $43/gb vs a $45k b200 alone)
The latter is standout for many things, including resource efficiency on smaller machines