← Back to post

Edit history

Most recent

Atm the meta for local inference is unified memory and “routed” local agents (multiple smaller role-specific agents in a trench coat)

The former is standout for cost efficiency (e.g., 4x RDMA 48gb Minis for a 192gb cluster @ $43/gb vs a $45k b200 alone)

The latter is standout for many things, including resource efficiency on smaller machines

Original

Atm the meta for local inference is unified memory and “routed” local agents (multiple smaller role-specific agents in a trench coat)

The former is standout for cost efficiency (e.g., 4x RDMA 48gb Minis for a 192gb cluster @ $0.43/gb vs a $45k b200 alone)

The latter is standout for many things, including resource efficiency on smaller machines