NEWTYPE · TECHNOLOGY · TECHNOLOGY — R&D

On any computing unit, training and inference, guaranteed

NEWTYPE keeps investing in R&D so that AI models can train and infer with guaranteed performance on any computing unit — from a single edge device to a heterogeneous distributed swarm that includes smartphones.

PATENT-PENDING

Edge AI Acceleration — patent-pending technology

“Apparatus and Method for Accelerating Artificial Intelligence Using Orthogonal Transform-Based Rotated Space Alignment” — an acceleration architecture for running both inference and training of large models on edge-device NPUs. A high-precision adapter in the same rotated space restores the accuracy lost to low-bit quantization.

  • Low-precision (INT4·FP4·NF4) base weights + high-precision (FP16·BF16·FP32) low-rank adapter — compensating quantization loss in the same rotated space
  • Fast Walsh–Hadamard Transform (FWHT) orthogonal transform — butterfly circuits doing only adds and subtracts, no floating-point multiplication
  • Part of the orthogonal transform is pre-merged into adjacent layers’ weights — hardware applies only the remaining transform to input and intermediate data
  • The adapter lives in on-chip memory, separated from the base-weight region — keeping only the update target in near memory
  • Intermediate values accumulate in a format of higher precision than the base — preventing error build-up from low-precision math
  • In backprop, the same (or inverse) transform maps error gradients into the rotated space — reusing the forward butterfly circuits to support on-device adapter updates (training)

How it works — 4 steps

  1. [1] Offline preprocessing

    Rotate → quantize

    In an external computing environment, input/output orthogonal matrices are applied to the original weights, quantized to a low-precision format, and shipped to the edge device.

  2. [2] Rotated-space alignment

    FWHT input transform

    The orthogonal transform aligns input data into the same rotated space as the base weights — using adds and subtracts only.

  3. [3] Dual-path compute

    BASE ∥ ADAPTER

    From the rotated-space input, the base path (low precision) and adapter path (high precision) compute in parallel.

  4. [4] Merge · update

    Output + on-device training

    The two results merge and the output-side transform is applied. In backprop, the same circuits transform gradients to update the adapter.

─ Patent application filed 2026.06 (examination requested)

DISTRIBUTED INFERENCE

Distributed Inference Engine

Run 100B+ models no single machine can hold — across many GPUs, machines, and even smartphones. A GPU-less controller plans layer and expert placement; an RPC data plane executes it.

  • A backbone stage runs attention, KV, norms, the router, the shared expert, and the combine — the router is an authority that runs once on the backbone
  • Expert workers are pure (hidden, local_ids) → out functions — only selected experts are dispatched to their owner nodes, and the backbone combines the results
  • The planner reads GGUF metadata to estimate contiguous per-node layer placement, KV-cache and expert VRAM, and even expert-FFN offload to node RAM
  • Nodes that cannot provide the resource monitoring needed for a safe plan are blocked from loading — no placement without a plan

Pipeline-parallel layer split

Reads GGUF metadata to plan per-node layer placement and KV-cache/expert VRAM, while a GPU-less master conducts many workers.

Expert-granular sharding

Most of a 122B MoE’s weight is thousands of ~5MB experts. Sharding by expert instead of layer lets even a 4GB phone hold and compute a slice. Router authority stays with the backbone.

Cross-backend numerical equivalence

CUDA, ROCm, Adreno, and CPU are not bit-identical but numerically equivalent. Router-authority design turns discrete divergence into bounded continuous error — a mixed-hardware swarm emits identical tokens.

Mobile nodes in the loop

A smartphone participated end-to-end in real 122B inference — autonomously downloading its expert slice, serving it, and computing one layer’s experts every token. Output tokens matched the reference exactly.

122B
Model size verified on a heterogeneous swarm
measured
1.0
Cosine similarity across GPU backends (CUDA vs ROCm)
measured
0.99975
Cosine similarity, GPU vs CPU
measured