Pipeline-parallel layer split
Reads GGUF metadata to plan per-node layer placement and KV-cache/expert VRAM, while a GPU-less master conducts many workers.
NEWTYPE · TECHNOLOGY · TECHNOLOGY — R&D
NEWTYPE keeps investing in R&D so that AI models can train and infer with guaranteed performance on any computing unit — from a single edge device to a heterogeneous distributed swarm that includes smartphones.
PATENT-PENDING
“Apparatus and Method for Accelerating Artificial Intelligence Using Orthogonal Transform-Based Rotated Space Alignment” — an acceleration architecture for running both inference and training of large models on edge-device NPUs. A high-precision adapter in the same rotated space restores the accuracy lost to low-bit quantization.
In an external computing environment, input/output orthogonal matrices are applied to the original weights, quantized to a low-precision format, and shipped to the edge device.
The orthogonal transform aligns input data into the same rotated space as the base weights — using adds and subtracts only.
From the rotated-space input, the base path (low precision) and adapter path (high precision) compute in parallel.
The two results merge and the output-side transform is applied. In backprop, the same circuits transform gradients to update the adapter.
─ Patent application filed 2026.06 (examination requested)
DISTRIBUTED INFERENCE
Run 100B+ models no single machine can hold — across many GPUs, machines, and even smartphones. A GPU-less controller plans layer and expert placement; an RPC data plane executes it.
Reads GGUF metadata to plan per-node layer placement and KV-cache/expert VRAM, while a GPU-less master conducts many workers.
Most of a 122B MoE’s weight is thousands of ~5MB experts. Sharding by expert instead of layer lets even a 4GB phone hold and compute a slice. Router authority stays with the backbone.
CUDA, ROCm, Adreno, and CPU are not bit-identical but numerically equivalent. Router-authority design turns discrete divergence into bounded continuous error — a mixed-hardware swarm emits identical tokens.
A smartphone participated end-to-end in real 122B inference — autonomously downloading its expert slice, serving it, and computing one layer’s experts every token. Output tokens matched the reference exactly.