NVIDIA’s new 64GB DGX Spark retains the GB10 Grace Blackwell Superchip, DGX OS and NVIDIA AI software stack of the 128GB configuration. It targets developers who want a dedicated machine for local AI agents without starting with the platform’s larger memory option.
Buyers can expand later. NVIDIA says two 64GB systems can connect through NVIDIA Sync Cluster Assistant, combining their capacity into a 128GB memory pool for larger workloads. A networked cluster, however, is different from one machine containing twice the memory.
The October 2, 2026, forum announcement schedules availability for October 23 from Acer, Dell, Gigabyte, HP and MSI. It does not confirm that those systems are already broadly shipping.
For prospective buyers, the question goes beyond whether 64GB can run an AI model. That capacity needs to fit the model, context length and number of agents they intend to keep running.
A Smaller Memory Option Without a Different Software Platform
The 64GB configuration lowers the starting capacity while keeping the same development environment. According to NVIDIA, it uses the same GB10 chip, operating system and full AI software stack as the 128GB model.
The DGX Spark product page identifies both capacities as coherent unified system memory. Within each system, that memory serves the CPU and GPU. There is no conventional split between system memory and a separate graphics-card memory allocation.
For developers already building around NVIDIA software, continuity is the main attraction. The announcement lists support for NVIDIA Agent Toolkit, CUDA-X AI libraries, Nemotron models and frameworks including Ollama, vLLM and PyTorch with CUDA. Developers choosing the smaller configuration are intended to retain that environment without changing toolchains.
NVIDIA says a single 64GB unit supports models of up to 100 billion parameters. That vendor-stated capacity limit is no promise that every model below that size will run comfortably.
Model precision, runtime overhead, context length and concurrent requests all affect an agent’s memory requirements. A heavily compressed model that fits in memory may leave less room for long conversations or several agents than its parameter count suggests.
The right buyer is someone with a reasonably defined workload: a coding assistant, document-analysis agent or local application whose chosen model and operating requirements fit within 64GB. Developers who already know they need substantially more capacity should evaluate the larger configuration or a cluster from the outset.
Local Agents Can Run Separately From Your Everyday PC
NVIDIA’s examples center on persistent agents, with one computer running the model and another used to interact with it.
A coding or research agent could remain running on DGX Spark, ready to review code, analyze documents or perform multistep tasks. In another proposed workflow, Spark handles inference while an agent interface or creative application runs on a laptop or desktop.
This gives developers a dedicated local model service without making their main computer responsible for every inference request. The intended role extends beyond occasional chatbot experiments to agents that remain available between interactive sessions.
NVIDIA Sync supports that working pattern. Its documentation describes a Windows, Mac and Ubuntu system-tray utility for connecting to remote Linux machines and launching applications, containers and development tools. It also handles connection details such as port forwarding, allowing applications running on Spark to remain accessible from another computer.
Local inference can reduce reliance on cloud token-generation services, but it does not automatically make an entire agent private or offline. An agent that calls an external API, searches the web or sends documents to a hosted tool still crosses that boundary.
Evaluating privacy requires looking at the complete workflow: where the model runs, which tools it can call, what data those tools receive and which permissions it holds. Hardware location answers only part of that question.
Two Systems Add Capacity, but Distributed Software Still Matters
NVIDIA’s clustering pitch is straightforward: start with one 64GB machine, then add another when the workload needs more memory or compute.
Each DGX Spark includes ConnectX-7 networking. The announcement says two units can connect directly through their ConnectX-7 ports with a QSFP cable, using the 200GbE fabric. NVIDIA describes the resulting two-system configuration as a 128GB memory pool supporting models of up to 200 billion parameters.
Sync Cluster Assistant handles the cluster infrastructure. The Sync user guide describes a guided workflow that validates devices, applies ConnectX-7 network settings, checks link performance and configures SSH between nodes.
At a high level, the setup involves:
- Connecting the two systems through their ConnectX-7 interfaces with an appropriate QSFP cable.
- Adding the devices to NVIDIA Sync so the application can access them.
- Running Cluster Assistant to validate and configure the interconnect.
- Launching a workload with software that supports execution across the nodes.
The workload still needs multi-node support. NVIDIA’s documentation describes cluster setup as the first step toward enabling libraries and runtimes, including NCCL and MPI, to use the high-bandwidth network. Cluster Assistant reduces the networking and connection work; it does not guarantee that an application written for one machine automatically becomes a distributed application.
Unified memory inside each Spark also differs from pooled capacity across two Sparks. The former is a hardware memory arrangement within one system. The latter depends on software distributing work across separate machines connected by a network. Both provide useful capacity, with different communication costs and execution requirements.
An agent developer could use the cluster to fit a larger model, accommodate longer contexts or support concurrent requests. Those needs are not interchangeable. A workload limited by model capacity may benefit differently from one limited by response latency.
NVIDIA also announces a Sync Model Launcher for the end of October. The company says it will download and launch Qwen3.8 27B on a single system or cluster and configure OpenCode to use the model. This separate software capability is planned for later; buyers should not assume it is already available with the October 2 announcement.
The Performance Claim Does Not Settle the Buying Decision
NVIDIA reports up to 1.7 times the performance from two clustered 64GB systems compared with a single system in its Qwen3.8 27B test.
The vendor-reported result suggests clustering can provide a speed benefit as well as additional capacity. It does not establish performance across other models or agent workloads.
Buyers should also take “up to” literally: every request will not necessarily complete 1.7 times faster, and doubling the number of systems does not necessarily double performance. Distributed inference introduces communication between nodes, with results depending on how the runtime divides the workload.
Sources
- 64GB DGX Sparkblogs.nvidia.com
- October 2, 2026, forum announcementforums.developer.nvidia.com
- DGX Spark product pagenvidia.com
- NVIDIA Syncdocs.nvidia.com





