Solution · Edge AI Host

Edge AI Host
RK3576 / RK3588 + RK1828 dual-chip collaboration

Run LLMs and multimodal models on an edge-side host, while HMIs, sensors and controllers degrade into lightweight "thin clients". A single host can carry the entire AI workload of a factory, a hotel or a home — letting front-end devices stay thin and bringing AI to every scenario.

RK3588 / RK3576 RK1828 LLM NPU Qwen3 1.7B–4B @ 84–138 TPS Qwen3-Omni 86 TPS RKNN3 Toolkit Edge–cloud collaboration

BesTom does not just sell core boards — we deliver Edge AI Host + thin clients + edge–cloud collaboration, a complete AI-transformation solution that drops into any small commercial or industrial scenario. One hardware base covers the full chain: perception — inference — execution — feedback.

Why this matters

Why an Edge AI Host

Offloading inference from the front end to a single edge host is the most practical path to bringing AI to industrial, commercial, hospitality and home environments.

HMIs become thin and light

Touch, display, sensing and protocol handling move to a low-power controller such as the RK3506; inference runs on the edge host. Finished-unit BOM and thermal load drop in parallel.

Unified intelligence across scenarios

One factory, one hotel or one home shares a single host and a single model set; behaviour and knowledge are shared across every terminal.

Data stays on-premise

Voice, image and behaviour data are understood and decided locally; sensitive information never leaves the site — meeting factory and hotel compliance requirements.

Works offline

Local inference keeps running during network outages, so critical scenarios (industrial control, hotel check-in, incident response) never drop.

Clear edge–cloud split

1.7B–4B models run locally; 7B+ and long-context tasks use edge–cloud collaboration — capacity added on demand, balancing cost and experience.

Smooth upgrade path

Swapping the chip on the host upgrades the whole network of front-end devices at once; no need to redo front-end hardware.

Architecture

Three-layer Collaborative Stack

Main-controller sandbox + edge-side brain + edge–cloud collaboration. One hardware base covers the full chain: perception — inference — execution — feedback.

LayerKey chipResponsibilityTypical capability
Main-controller sandbox
Main Sandbox
RK3576 / RK3588
(CPU + GPU + NPU)
General compute, system scheduling, agent framework, perception input / execution output, private-data storage Agent OS / Claw framework · general development base · complex business orchestration
Edge-side brain
Edge Brain
RK1828
(LLM NPU)
Local inference of large-language / multimodal models, local skill execution, private-data processing 4B model, 32K context · local skill calls within 10K · simple tasks closed-loop locally
Edge–cloud
Edge–Cloud
Cloud large-model token plans Complex task planning, deep long-context reasoning, on-demand compute top-up 7B+ long context · cross-domain knowledge base · training / fine-tune feedback
Front-end thin clients
Thin Clients
RK3506 / RK2116 / RV1106B etc. Touch, voice I/O, sensor acquisition, electromechanical control, local pre-processing 86-box panels · massage-chair panels · IPC · industrial PLC panels · home central control
Performance

On-device Throughput

The RK1828 is purpose-built for large-language / multimodal inference, delivering 6–10× the speed of a CPU-only path. Figures below are measured at W4A16 / W8A8-3 quantization, 1024-token context, 128 in / 128 out.

PlatformModelQuantizationTTFT (ms)Prefill (tok/s)TPS (tok/s)Memory (MB)Bandwidth (MB/s)
RK182XQwen3.5-2BW4A16362.7352.941.13151082 378
Qwen3-1.7BW4A1653.62387.5137.98133316 150
Qwen3-4BW4A16107.81187.584.102696220 029
RK3588Qwen3.5-2BW8A8-3923.3138.612.76210927 442
Qwen3-1.7BW8A8-3429.3298.213.88193426 904
Qwen3-4BW8A8-31012.0126.56.75418428 817
RK3576Qwen3.5-2BW4A161726.774.19.41146211 315
Qwen3-1.7BW4A16807.7158.611.32127111 620
Qwen3-4BW4A161733.173.85.57246012 386
Source: Rockchip RK182X SDK 1.1.0 Preview platform performance summary. Actual rates vary with context length, batch size and system load.
Model library

Supported Model Zoo

The RK1828 + RKNN3 toolchain covers language, voice, vision, multimodal and omni-modal model families.

CategoryRepresentative models
LLMQwen2.5-0.5B/1.5B/3B/7B · Qwen3-0.6B/1.7B/4B/8B · GLM Edge · Hunyuan-MT1.5 · Youtu-LLM
Voice ASR / TTSSenseVoice · Qwen3-ASR / TTS · Whisper · VITS
Multimodal / VLMQwen3.5-9B · Gemma4-12B · Qwen2.5-VL 3B/7B · Qwen3-VL 2B/4B · Qwen3.5 2B/4B · MiniCPM-V-4 · PaddleOCR
Omni-modalQwen3-Omni · Qwen3.5-Omni · Gemma4-E2B / E4B
CategoryRepresentative models
Vision ViT / CNNSigLIP 1/2 · EVA02 · DINOv2/v3 · CLIP ViT-B/L · Yolo V5/V6/V8/World/26 · MobileNet · ResNet
Depth / localizationDepth-Anything-V2-small · LocateAnything-3B
EmbeddingM3E-small · Albert-base-v2
RerankerQwen3-0.6B-Reranker · Qwen3-4B-Reranker
Text-to-imageDream-Lite
Toolkit

RKNN3 Toolkit

RKNN3 turns "model → on-device executable" into a single command, with custom operators and very low-memory conversion.

Upgrade 1 · Custom CPU operators

Customers can develop their own inference operators for bespoke business scenarios, extending model and heterogeneous-acceleration freedom.

Upgrade 2 · Usability optimization

Removes third-party whl package version limits, simplifies environment deployment, and greatly lowers the GPU memory needed for model conversion; context length is quick to change, and assessment / deployment modes switch with one click.

Quantization strategy

W4A16 / W8A8-3 and other multi-tier quantization levels switch freely by accuracy and speed.

Cross-platform export

One model packages to both the RK1828 NPU and the RK3588 / RK3576 CPU/GPU/NPU, flexibly matching host-chip combinations.

Newly added

Latest Additions

Newest RKNN3 adaptations: omni-modal, multimodal and vision specialists.

CategoryModelThroughputNote
Omni-modalQwen3-Omni86 TPSLocal omni-modal interaction
Qwen3.5-Omni62 TPSUpgraded omni-modal version
Gemma4-12B30 TPSLarge omni-modal model
Multimodal VLMQwen3.5-4B50 TPSVision-language 4B
Qwen3.5-9B35 TPSVision-language 9B
Vision specialistDream-Lite3.1 s / imageText-to-image
LocateAnything-3B102 TPSVision localization & referring
Application

Where it deploys

The same "Edge Host + Thin Clients" base becomes a general-purpose AI-transformation engine — it drops into any small commercial or industrial scenario, unifying distributed devices and unifying intelligence. Typical landing directions below.

Small commercial · hotel / storefront / F&B

Guest rooms / front desk / ordering / queuing unified on one edge host; thin HMI and IoT devices act as thin clients. Virtual-human reception, voice temperature control, people counting — all in one place.

Small commercial · clinic / retail / warehouse

Local voice entry, goods & inventory vision counting, body-temperature screening — data never leaves the store, privacy compliant.

Industrial · factory / production line / building

One host per plant; PLC panels / HMIs / sensors all become thin clients; inference is centralized and upgrades only touch the host.

Industrial · predictive maintenance

Vibration / temperature / sound multimodal anomaly detection locally, alerting without going to cloud — works offline.

Smart home

One host per home; HMI uses thin panels; whole-home voice / vision understanding runs as a local closed loop.

Office meeting / vehicle cabin / education

Meeting transcription, offline in-vehicle multimodal, on-campus local tutor — no drop during outages.

Reference deployment

Reference Solution · 200-room Hotel AI Transformation

Using a 200-room hotel as an example, see how "Edge Host + Thin Clients" completes an AI transformation in one pass.

LocationDeviceRoleKey capability
Each guest room × 200RK3506 86-box / wall panelThin-client HMITouch + voice I/O, lighting control, HVAC, curtain, scenes, energy reporting; no large model on device
Front desk / server room × 1RK3576 + RK1828 edge hostBrainVirtual-human reception · RAG knowledge base · voice temperature control · entrance people counting · scenario-wide agent
Public area / floorRK2116 audio moduleSpatial audioMulti-mic array pickup · spatial audio · far-field wake · cross-room voice routing
Entrance / corridorRV1106B / RV1126B IPCEdge visionFace-recognition access control · people counting · abnormal-behaviour detection
Economics: guest-room HMI only needs a low-power controller; BOM and thermal load drop, while the whole hotel's compute concentrates on one edge host.
Intelligence: virtual-human reception + RAG + multimodal voice/vision unified on the edge host; upgrade only replaces host hardware.
Compliance: voice, image and behaviour data are processed locally; sensitive information never leaves the hotel.
Why Bestom

Why BesTom

The hard part of "Edge Host + Thin Clients" is not the chip itself, but the edge–cloud split, protocol adaptation and on-site delivery. BesTom has mature end-to-end capability across the full Rockchip line.

Rockchip full-line IDH

RK3588 / RK3576 / RK3506 / RV1106B / RK1828 / RK2116 integrated in one place — no battery of multi-supplier assembly.

Model engineering

Closed loop from Qwen3 1.7B–4B to Qwen3-Omni quantization, conversion and on-device deployment.

Thin-client experience

Validated "interaction-only" thin-client design on the 86-box, massage-chair panel and industrial HMI.

Edge–cloud delivery

Local inference + large-model token-plan hybrid architecture, with graceful degradation and upgrade paths for the whole network.

Protocol & on-site

Multi-protocol adaptation for industrial / hotel / home; integrated mass-production delivery, DFM, thermal simulation and EMC.

Compliance & certification partners

Access to CE / FCC / RoHS and medical-related registration guidance resources, shortening time to market.

Related

Related Resources

Massage-chair HMI Solution

Front-end control panel RK3506 design and core-mechanism interfacing.

Rockchip SoM

RK3588 / RK3576 / RK3506 / RK1828 SoM and development-board details.

Inquire

Bring your scenario and model needs; NRE, pilot run and certificationcadence can be assessed together.