# Vendor Deployment Guide
Source: https://docs.rminte.com/essentials/deploy
Comprehensive guide for vendors to deploy RM-01 Portable AI Supercomputer solutions for enterprise customers, achieving a true “plug-and-play” AI experience.
## Overview
This guide covers the complete deployment process from preparation to after-sales support:
Model preparation, environment assessment, and device verification
Installation, initialization, and network configuration
Ongoing support, maintenance, and customer service
## Pre-deployment Preparation
### Model and Application Preparation
Prepare suitable models and applications based on enterprise requirements:
Includes common open-source models such as:
* DeepSeek
* Zhipu
* Qwen
* Other mainstream models
Standard models are pre-configured and ready for immediate deployment across various enterprise use cases.
Vertical domain models tailored for specific industries:
* Healthcare and medical AI
* Financial services models
* Manufacturing optimization
* Legal and compliance models
Industry-specific models often provide better performance for specialized enterprise workflows.
Custom trained models specific to enterprise needs:
* Convert models to RM-01 compatible formats
* Store in specified directory on CFexpress card
* Ensure proper model validation and testing
Custom models require thorough validation to ensure compatibility with RM-01's hardware specifications.
### Enterprise Environment Assessment
Conduct thorough environment assessment before deployment:
Assess the feasibility of connecting RM-01 to the enterprise internal network:
* Evaluate network topology and security requirements
* Identify firewall and proxy configurations
* Plan IP addressing and network segmentation
* Test network bandwidth and latency requirements
Document network requirements and obtain necessary approvals from IT security team.
Ensure stable power supply at the deployment location:
* Verify power outlet availability and specifications
* Check for uninterruptible power supply (UPS) requirements
* Assess power consumption impact on facility
* Plan for power redundancy if needed
RM-01 requires 100W maximum power consumption with 60W TDP for optimal performance.
Assess physical security of device placement:
* Evaluate physical access controls
* Review environmental conditions (temperature, humidity)
* Plan for device mounting and cable management
* Assess theft prevention measures
Ensure deployment location meets enterprise security standards and environmental requirements.
Understand existing IT infrastructure:
* Review current AI/ML infrastructure
* Assess integration with existing systems
* Plan for user access management
* Evaluate monitoring and logging requirements
### Device and Accessories Verification
Before delivery, ensure all components are complete and functional:
**Core Components:**
* RM-01 Host device
* CFexpress Type B Storage Card (pre-installed with models and applications)
* MicroSD Card for system storage
* Quick Start Guide with setup instructions
* Warranty Card and documentation
Verify all components are present and undamaged before shipment.
**Essential Accessories:**
* USB-C Power Adapter (PD3.1, up to 140W)
* USB-C to Ethernet Adapter for network connectivity
* Additional cables as needed
Keep spare accessories available for immediate replacement if needed.
## Hardware Deployment Process
### Device Installation
Prepare the installation location:
1. Choose optimal placement (desktop or dedicated cabinet)
2. Ensure adequate ventilation around the device
3. Verify power and network connectivity
4. Prepare cable management solutions
Ensure the device is placed in a well-ventilated environment, avoiding obstruction of heat dissipation vents.
Install the RM-01 device:
1. Insert CFexpress storage card with pre-installed models and applications
2. Insert MicroSD card for system storage
3. Connect USB-C power adapter
4. Connect USB-C to Ethernet adapter if wired network is required
Handle storage cards carefully to avoid damage to connectors and data corruption.
### System Initialization
Power on and initialize the system:
1. Press the power button to start the device
2. Monitor initialization process via connected laptop (USB-C)
3. Wait for automatic model and application loading (\~5 minutes)
4. Verify system reaches ready state
During initialization, you can connect a laptop via USB-C port to monitor device status and troubleshoot if needed.
Confirm successful initialization:
* Check system status indicators
* Verify model loading completion
* Test basic device responsiveness
* Validate storage card recognition
System should display ready status after approximately 5 minutes of initialization.
### Network Configuration
Configure network connectivity based on enterprise requirements:
Configure Ethernet connection:
```bash theme={null}
# Static IP Configuration
sudo ip addr add 192.168.1.100/24 dev eth0
sudo ip route add default via 192.168.1.1
# DNS Configuration
echo "nameserver 8.8.8.8" | sudo tee /etc/resolv.conf
```
Verify network connectivity with ping test to external servers.
Enable automatic IP assignment:
```bash theme={null}
# Enable DHCP client
sudo dhclient eth0
# Verify IP assignment
ip addr show eth0
```
DHCP is recommended for simplified network management in most enterprise environments.
## Verification and Testing
### Comprehensive Testing Protocol
Verify core functionality:
**Model Performance:**
* Test model loading and execution
* Verify inference speed and accuracy
* Check resource utilization
**Application Testing:**
* Test basic application functionality
* Verify user interface responsiveness
* Check integration with enterprise systems
Document test results and compare against expected performance benchmarks.
Conduct performance evaluation:
**Metrics to Test:**
* Model inference latency
* Concurrent processing capability
* Memory and storage utilization
* Network throughput
```bash theme={null}
# Example performance test command
./benchmark_tool --model deepseek --iterations 100 --concurrent 4
```
Use standardized benchmarking tools to ensure consistent performance evaluation across deployments.
Verify security measures:
* Test user authentication and authorization
* Verify data encryption at rest and in transit
* Check network access controls
* Validate audit logging functionality
Ensure all security configurations meet enterprise compliance requirements before deployment approval.
## Training and Delivery
### Administrator Training Program
Train administrators on basic device management:
**Core Topics:**
* Device startup and shutdown procedures
* Hardware maintenance and troubleshooting
* Storage card management
* Network configuration updates
Demonstrate system management capabilities:
* Web-based management console
* System monitoring and alerts
* User account management
* Configuration backup and restore
Record training sessions for future reference and onboarding new administrators.
Cover advanced management topics:
* Model updates and version control
* Application deployment procedures
* Performance optimization techniques
* Troubleshooting common issues
### End User Training
Train users on fundamental operations:
* Application access and login procedures
* Basic AI model interaction
* File upload and processing
* Result interpretation and export
Customize training content based on specific departmental use cases for maximum effectiveness.
Cover sophisticated functionality:
* Multi-model workflow creation
* Custom prompt engineering
* Batch processing operations
* Integration with external tools
Provide hands-on practice sessions with real enterprise data to enhance learning effectiveness.
### Documentation Delivery
Provide comprehensive documentation package:
Hardware specifications and basic operation guide
System management and troubleshooting procedures
Application usage steps and best practices
Common issues and resolution procedures
All documentation is provided in both electronic and printed formats, with the latest versions available through the management interface.
## After-sales Support
### Support Service Levels
**Included Services:**
* 12 months of remote support
* Business hours response (9 AM - 6 PM)
* Email and phone support
* System health monitoring
**Response Times:**
* 4-hour response during business hours
* 24-hour resolution for standard issues
Standard support covers most enterprise deployment needs with comprehensive coverage.
**Enhanced Services:**
* 24/7 technical support availability
* 2-hour response time guarantee
* Priority issue resolution
* Dedicated support engineer
**Additional Features:**
* Proactive system monitoring
* Quarterly performance reviews
* Priority access to updates
Premium support is recommended for mission-critical deployments and high-availability requirements.
**Comprehensive Services:**
* 30-minute emergency response
* On-site support when needed
* Custom model development assistance
* Integration consulting services
**Value-Added Services:**
* Performance optimization consulting
* Custom application development
* Staff augmentation services
Enterprise support provides the highest level of service for complex deployments and specialized requirements.
### Support Contact Information
**Phone:** 158-8200-8185
**Hours:** Monday - Friday, 9:00 AM - 9:00 PM
Direct access to technical support engineers
**Email:** [support@rminte.com](mailto:support@rminte.com)
**Response:** 24-hour guarantee
Detailed technical inquiries and documentation
**Website:** rminte.com
**Chat:** Real-time during business hours
Instant support and resource access
## Frequently Asked Questions
No. RM-01 is designed to operate completely offline, ensuring data security and independence from internet connectivity. All processing occurs locally on the device.
RM-01 has a maximum power consumption of 100W with a TDP of 60W, saving 98.6% energy compared to traditional AI servers. Annual electricity cost for 24/7 operation is approximately \$480.
RM-01 supports models ranging from 0.5B to 235B parameters (GPTQ Int4 quantization), covering all mainstream open-source models including DeepSeek, Qwen, and Llama variants.
RM-01 uses hardware-level asymmetric encryption with all data processed locally. No data is ever uploaded to the cloud, ensuring complete enterprise data sovereignty and security.
Compared to traditional AI deployment, RM-01 saves 80% in initial investment, reduces operational costs by 98%, and achieves 99% total cost of ownership savings over 3 years.
Yes, we offer a 7-day free trial service for enterprise customers with professional technical staff providing on-site support and evaluation assistance.
## Technical Support Resources
Comprehensive technical documentation and API references
Technical forums and best practice sharing
Discover and deploy professional AI applications
Video tutorials and certification programs
***
# RM-01 Developer Guide
Source: https://docs.rminte.com/essentials/develop
Comprehensive guide to RM-01 system architecture, module configuration, network setup and model deployment
## Overview
This guide provides developers with complete technical documentation for the RM-01 portable supercomputer, covering system architecture, network configuration, model deployment and other core content:
Configure network connections for data interaction between device and host
Understand how the inference module, application module and management chip work together
Master AI model deployment, configuration and optimization methods
**Read Before Use**
RM-01 consists of an **Inference Module**, an **Application Module**, and an **Encryption and Management Chip** (hereinafter referred to as the Management Module), interconnected via an **onboard Ethernet switch chip**, forming an internal LAN subnet. When a user connects to a host (e.g., PC, smartphone, iPad) via the **USB Type-C** interface, RM-01 will virtualize an Ethernet interface for the host through USB Ethernet functionality. The host will then obtain an IP address and automatically join the subnet for data interaction.
After the device is powered on and connected to the host via the **USB Type-C** interface, the system will automatically configure the local network subnet. The **user host** will be assigned a static IP address `10.10.99.100`, and the **Out-of-Band Management Chip** will have a static IP address of `10.10.99.97`. The **Inference Module** (IP: `10.10.99.98`) and the **Application Module** (IP: `10.10.99.99`)—both deploy independent SSH services, allowing users to access them directly via standard SSH clients (e.g., OpenSSH, PuTTY). The Management Module, however, requires access via a serial port tool.
## Network Configuration
### About the Out-of-Band Management Chip
**Network Configuration**
* IP Address: `10.10.99.97`
* Access Method: Web browser
In addition to handling certain encryption tasks, the Out-of-Band Management Chip also hosts the **RM-01's real-time system performance monitoring dashboard** — **RobOS**.
You can access `10.10.99.97` through a web browser to monitor the **connection status** and **operational status** of each module in real-time.
### How to Provide Internet Access to RM-01 from the Host (Using macOS as an Example)
After connecting the **user host** via USB Type-C, RM-01 will appear in the network interface list as:
* **`AX88179A`** (Developer Version)
* **`RMinte RM-01`** (Commercial Release Version)
Open **System Settings**
Go to **Network** → **Sharing**
Enable **Internet Sharing**
Click the **"i" icon** next to the sharing settings to enter the configuration interface:
* Set **"Share your connection from"** to: **Wi-Fi**
* In **"To computers using"**, select: **AX88179A** or **RMinte RM-01** (depending on the device model)
Click **Done**
Return to the **Network** settings page and manually configure the RM-01 network interface:
* **IP Address**: `10.10.99.100`
* **Subnet Mask**: `255.255.255.0`
* **Router**: `10.10.99.100` (i.e., the host's own IP)
This configuration sets the host as a gateway, providing NAT network access for RM-01. The default gateway and DNS for RM-01 are automatically assigned by the host via DHCP. Manually setting the IP ensures that it remains within the `10.10.99.0/24` subnet, consistent with the device's internal service communication.
## System Architecture
### About the CFexpress Type-B Storage Card
The **CFexpress Type-B** storage card is one of the core components of the RM-01 device, responsible for system boot, deployment of the inference framework, and key functions such as ISV/SV software distribution and authorization authentication.
The storage card is divided into three independent partitions:
**System Partition**
The operating system and core runtime environment of the Inference Module are installed in this partition.
Users or developers are strictly prohibited from accessing, modifying, or deleting the contents of this partition. Any unauthorized changes may cause the Inference Module to fail to boot or render inference functions inoperable, and any resulting hardware or software damage is not covered by any warranty services.
**Application Partition**
This partition is used to temporarily store `Docker` image files submitted by users or developers. After the image is written to `rm01app`, the RM-01 system will automatically migrate it to the device's built-in **NVMe SSD** storage and complete containerized deployment.
Do not directly run or modify application files in this partition.
**Model Partition**
Dedicated to storing large-scale AI models (e.g., LLMs, multimodal models, etc.) loaded by users or developers.
For details on model formats, size limitations, loading procedures, and compatibility requirements, refer to the "Model Deployment" section below.
### About the Application Module
**Network Configuration**
* IP Address: `10.10.99.99`
* Port Range: `59000-59299`
#### Application Module Hardware Specifications
```text theme={null}
Processor: Intel Core i3-N305 (8 cores, 8 threads, base frequency 1.8 GHz, max turbo frequency 3.8 GHz)
Memory: 16 GB / 24 GB LPDDR5-4800MT/s (onboard, non-expandable)
Storage: 512 GB / 1 TB / 2 TB (optional) NVMe SSD
```
#### Application Module SSH Access Credentials
```bash SSH Login Command theme={null}
ssh rm01@10.10.99.99
# Default Username: rm01
# Default Password: rm01 (factory preset, for initial login only)
```
```bash Change Password theme={null}
# Execute immediately after first login
passwd
```
**Security Notice**
To ensure system security, immediately use the `passwd` command to change the default password after the first SSH login. The default password is only for initial configuration and must not be used in production or deployment environments.
#### Pre-installed Software
The Application Module has **Open WebUI** pre-installed on port `80` to facilitate simple model debugging and conversational work.
You can access Open WebUI by navigating to `10.10.99.99` in your web browser.
### About the Inference Module
**Network Configuration**
* IP Address: `10.10.99.98`
* Service Port Range: `58000–58999`
The **Inference Module** is the core computing unit of RM-01, supporting various high-performance AI inference configurations. Users can select the appropriate model based on model scale and performance requirements.
#### Hardware Configuration Options
| Memory | Memory Bandwidth | Compute Power | Tensor Core Count |
| :----: | :--------------: | :----------------: | :---------------: |
| 32 GB | 204.8 GB/s | 200 TOPS (INT8) | 56 |
| 64 GB | 204.8 GB/s | 275 TOPS (INT8) | 64 |
| 64 GB | 273 GB/s | 1,200 TFLOPS (FP4) | 64 |
| 128 GB | 273 GB/s | 2,070 TFLOPS (FP4) | 96 |
#### Pre-installed Inference Frameworks
RM-01 comes pre-installed with the following two inference frameworks on the **CFexpress Type-B** storage card, both running on the Inference Module:
* **Status**: Automatically starts
* **Default Port**: 58000
* **Function**: Provides OpenAI-compatible API interfaces
* **Supported Requests**: Standard POST `/v1/chat/completions` etc.
* **Status**: Requires manual startup
* **Function**: Text embedding services
#### API Access Method
After successfully loading a model, the **vLLM** inference service can be accessed via the following address:
```bash theme={null}
http://10.10.99.98:58000/v1/chat/completions
```
Supports direct calls using standard OpenAI clients (e.g., openai-python, curl, Postman).
**Security Notice**
To ensure system security and stability, the Inference Module does not provide SSH access permissions. Users and developers cannot directly log in or interactively operate the underlying operating system of this module.
Any attempts to bypass security policies or directly access the Inference Module may result in system anomalies, data corruption, or service interruptions, which are not covered by warranty services.
## Model Deployment
### About Models
RM-01 supports inference for various AI models, including but not limited to:
Large Language Models
Multimodal Models
Vision-Language Models
Text Embedding Models
Reranker Models
All model files must be stored on the device's built-in **CFexpress Type-B** storage card, and users need to use a compatible **CFexpress Type-B** card reader to upload, manage, and update models on the host side.
When the **CFexpress Type-B** storage card is connected to RM-01, the system mounts it as a read-only data volume named `models` at the path `/home/rm01/models`. Its standard file structure is as follows:
```bash theme={null}
models/
├── auto/ # Directory for automatic model loading (production-grade deployment)
│ ├── embedding/ # Embedding models (not automatically loaded, see below)
│ ├── llm/ # Large language models (weight files stored directly, see below)
│ └── reranker/ # Reranker models (not automatically loaded, see below)
└── dev/ # Developer-defined model directory (higher priority)
├── embedding/
├── llm/
└── reranker/
```
* The `auto/` directory is used for **lightweight, standardized model deployment**, automatically recognized by the system
* The `dev/` directory is used for **fine-grained control of model behavior by developers**, with higher priority than `auto/`. The system will ignore models in `auto/` if `dev/` is used
### Deployment Mode Selection
Simplified mode suitable for quick verification and standardized deployment.
Place the **complete weight files** of the model (e.g., `.safetensors`, `.bin`, `.pt`, `.awq`, etc.) **directly in the `auto/llm/` directory**, **without nesting in subfolders**.
```text ❌ Incorrect Example theme={null}
auto/llm/Qwen3-30B-A3B-Instruct-2507-AWQ/model.safetensors
```
```text ✅ Correct Example theme={null}
auto/llm/model-001-of-006.safetensors
auto/llm/config.json
auto/llm/tokenizer.json
auto/llm/vocab.json
```
* Upon device startup, the system scans the `auto/llm/` directory and automatically loads models in compatible formats
* **Automatic loading is not supported for embedding or reranker models**, only LLMs
* After loading, the model enables **basic inference capabilities by default** and **does not enable the following advanced features**:
* Speculative Decoding
* Prefix Caching
* Chunked Prefill
* The maximum context length (`max_model_len`) is restricted to the system's safe threshold (typically ≤ 8192 tokens)
* **Limited performance optimization**: To ensure system stability and multitasking concurrency, models in automatic mode use a conservative memory allocation strategy (`gpu_memory_utilization` ≤ 0.8)
**Important Note**
Automatic mode is suitable for **quickly verifying model compatibility** or **standardized deployment scenarios**, but it is **not suitable for high-performance inference in production environments**. For full performance, use **manual mode (dev)**.
Developer mode, providing complete configuration control and performance optimization capabilities.
**Manual mode has higher priority than automatic mode**. When any `.yaml` configuration file exists in the `embedding`, `llm`, or `reranker` directories under `dev/`, the system will **completely ignore all models in `auto/`**.
#### Configuration File Structure
The `embedding`, `llm`, and `reranker` directories under `dev/` must contain the following three YAML configuration files to start the corresponding model services:
| File Name | Purpose |
| -------------------- | ---------------------------------- |
| `embedding_run.yaml` | Start text embedding model service |
| `llm_run.yaml` | Start large language model service |
| `reranker_run.yaml` | Start reranker model service |
All files must be located in the corresponding folders under `dev/`, and **file names must match exactly**, case-sensitive.
#### Example Configuration File
```yaml llm_run.yaml theme={null}
model: /home/rm01/models/dev/llm/Qwen3-30B-A3B-Instruct-2507-AWQ
port: 58000
gpu_memory_utilization: 0.85
max_model_len: 24576
served_model_name: RM-01 LLM
enable_prefix_caching: true
enable_chunked_prefill: true
max_num_batched_tokens: 512
block_size: 16
tensor_parallel_size: 1
dtype: auto
```
**Configuration Notes**
* The `model` path must be an **absolute path** pointing to the model file directory (not a compressed file)
* The `port` must be within the `58000–58999` range and must not conflict with other services
* The `gpu_memory_utilization` is recommended to be set between 0.7–0.9 to maximize throughput
* All parameters follow the [vLLM official documentation](https://docs.vllm.ai/en/latest/index.html). Ensure compatibility with the vLLM version currently used on the **CFexpress Type-B** storage card
#### Deployment Steps
Copy the model files (complete directory) to `dev/llm/` (or `embedding/`, `reranker/`)
Create and place the corresponding `.yaml` configuration file
Insert the **CFexpress Type-B** storage card and restart RM-01
The system will automatically load the configuration in `dev/` and start the corresponding service
Once the service is started, it can be accessed via interfaces like `http://10.10.99.98:58000/v1/chat/completions`
**Recommendation**: For initial deployment, use the `--load-format auto` and `--dtype auto` options provided by vLLM to automatically adapt to the model format.
### Security and Maintenance Notes
* **Prohibited SSH Login to Inference Module**: All model management must be done via the **CFexpress Type-B** storage card
* **Model Files Must Be Raw Weights**: Do not use compressed files (.zip/.tar.gz), encrypted packages, or non-standard formats
* **File Permissions**: All model files must be readable (`chmod 644`), and directories must be executable (`chmod 755`)
* **Version Control**: It is recommended to use Git or file naming conventions (e.g., `Qwen3-30B-A3B-Instruct-v1.2-20250930`) to manage model versions
* **Backup Recommendation**: Back up the `dev/` and `auto/` directories before updating models to avoid configuration loss
### Mode Selection Recommendations
| Scenario | Recommended Mode | Description |
| ------------------------------------------- | -------------------------------------------- | ---------------------------------------- |
| Quickly verify model compatibility | Automatic Mode (auto) | No configuration required, plug and play |
| High-performance inference in production | Manual Mode (dev) + fine-tuned configuration | Full performance optimization |
| Multi-model parallel deployment | Manual Mode (dev) + multiple `.yaml` files | Flexible service orchestration |
| Development debugging, prototype validation | Manual Mode (dev) | Complete control |
## Technical Support
Complete API reference and technical documentation
Sample code and open-source tools
Developer community and technical discussions
Professional technical support services
***
# Enterprise Quickstart
Source: https://docs.rminte.com/essentials/quickstart
Get your team up and running with the RM-01 AI supercomputer quickly with this comprehensive enterprise guide
## Device Overview
The RM-01 AI supercomputer is a portable AI solution that combines high performance, security, and ease of use, specifically designed for enterprise-level AI applications. It operates offline with local computation, significantly reducing data security risks and operational costs.
### System Architecture
RM-01 consists of three core modules interconnected through a built-in Ethernet switch chip, forming an independent internal network:
**IP Address**: `10.10.99.98`
Handles AI model inference computation, providing high-performance computing power
**IP Address**: `10.10.99.99`
Runs pre-installed applications and business systems, such as Open WebUI
**IP Address**: `10.10.99.97`
Provides RobOS real-time monitoring and device management functions
When you connect RM-01 to your host device (computer, phone, or tablet) via USB Type-C, the host will automatically be assigned the static IP address `10.10.99.100` and join RM-01's internal network.
### Key Features
* Supports 300+ mainstream open-source models
* Up to 2070 TFLOPS computing power
* 128GB LPDDR5 high-speed memory
* Hardware-level asymmetric encryption protection
* Data processed locally, never transmitted externally
* Anti-tamper physical security design
* Deployment setup completed within 5 minutes
* Plug-and-play, no professional IT support needed
* Maximum power consumption only 140W (TDP 60W)
### Device Interfaces
* **Power Button**: Short press to turn on/off, long press for 5 seconds to force shutdown
* **Status Indicator**: White (Running), Blue (Starting), Red (Error)
* **CFexpress Slot**: For model and application storage, supports hot-swapping
* **MicroSD Slot**: For logo customization and key storage
* **Expansion Port**: For connecting external expansion devices (supports multimodal input)
* **USB-C Power Port**: Connect to power adapter (supports PD3.1)
* **USB-C Network Port**: Connect to personal devices or enterprise switches
* **Ventilation**: Ensure good ventilation, do not block
## Quick Start
### First-time Setup
Confirm the package contains the following items:
* RM-01 Main Unit
* CFexpress Card with pre-installed models and applications
* MicroSD Card
* Warranty Card
* Quick Start Guide
* USB-C Power Adapter
* USB-C to Ethernet Adapter
* Connect the USB-C power adapter to the RM-01
* Connect the power adapter to a power outlet
* Press the power button
* Wait for the device to boot (about 1-2 minutes)
* A steady white indicator light means the device is running normally
Use a USB Type-C cable to connect RM-01 to your host device (computer, phone, or tablet). After the device is powered on, it will automatically configure the internal network, and your host will be assigned the static IP address: `10.10.99.100`
After successful connection, your host's network interface will display as **AX88179A** or **RMinte RM-01** (depending on version)
If you need RM-01 to access the internet, you can share the network through your host device. For macOS example:
1. Open **System Settings** > **Network** > **Sharing**
2. Enable **Internet Sharing**, select **Wi-Fi** as the sharing source
3. In "Share to devices", check the RM-01 network interface
4. Manually set the RM-01 network interface parameters:
* **IP Address**: `10.10.99.100`
* **Subnet Mask**: `255.255.255.0`
* **Router**: `10.10.99.100`
5. Click **Done**
This configuration makes the host act as a gateway, and RM-01 will automatically obtain the default gateway and DNS. If internet access is not needed, you can skip this step.
### Accessing the Management Interface
RM-01 provides multiple management entry points, each corresponding to different functional modules:
Access the real-time monitoring dashboard of the management module:
```
http://10.10.99.97
```
View the connection status and operation of each RM-01 module here, keeping track of device health in real-time.
Access the pre-installed Open WebUI on the application module:
```
http://10.10.99.99
```
**Default login credentials**:
* Username: `rm01`
* Password: `rm01`
After first login, please use SSH to log in (address: `10.10.99.99`) and immediately change the password to ensure security.
Directly access the API service of the inference module:
```
http://10.10.99.98:58000/v1/chat/completions
```
Supports standard OpenAI client calls, using tools like `openai-python`, Postman, or curl.
If you have purchased or deployed other enterprise applications, they typically run on the application module (`10.10.99.99`). Please contact your service provider for specific port numbers and access methods.
## Model Management
### Loading Open-Source Models (Automatic Mode)
RM-01 supports quick loading of open-source large language models (LLM), allowing you to easily verify model compatibility.
Using a **CFexpress Type-B** card reader, place the complete model weight files into the `auto/llm/` directory of the storage card.
Supported formats include: `.safetensors`, `.bin`, `.pt`, `.awq`, etc.
**Correct example**:
```
auto/llm/model-001-of-006.safetensors
auto/llm/config.json
auto/llm/tokenizer.json
auto/llm/vocab.json
```
Do not use subfolders! Files must be placed directly in the `auto/llm/` directory, not in subdirectories.
**Incorrect example**:
```
auto/llm/Qwen3-30B-A3B-Instruct-2507-AWQ/model.safetensors
```
Insert the **CFexpress Type-B** storage card into RM-01 and restart the device. The system will automatically scan the `auto/llm/` directory and load compatible models.
After the model is loaded, you can access the inference service through the inference service endpoint (`http://10.10.99.98:58000/v1/chat/completions`).
**Automatic Mode Limitations**:
* Only supports large language models (LLM), does not support embedding models or reranker models
* Suitable for quick verification of model compatibility, performance is limited (maximum context length ≤ 8192 tokens, memory usage rate ≤ 0.8)
* For higher performance or multi-model parallel deployment, please contact your service provider for **manual mode (dev)** technical support
## Using AI Applications
### View and Launch Applications
1. Log in to the management interface
2. Click the "App Library" tab
3. Browse the list of installed AI applications
Click on the application name to view details, including functionality introduction, usage scenarios, and required resources
1. Click the "Launch" button
2. The application will open in a new tab
3. Follow the application prompts to operate
## Data Management
### Data Import
1. Find the "Upload" or "Import" button in the application
2. Select the file to upload
3. Wait for the upload to complete
1. Prepare the files to be imported
2. Use the provided batch import tool
3. Select the folder or multiple files
4. Start the import process
### Data Security Guarantees
RM-01 uses multiple security measures to protect your data
All data is processed within the device, never uploaded to the cloud
Uses asymmetric encryption technology to protect stored data
Dust and moisture-proof design, prevents physical disassembly
Granular user permission management system
### Security Usage Notes
**The inference module does not support SSH login**. All model management must be completed through the **CFexpress Type-B** storage card. This design ensures the stability and security of the inference service.
Do not attempt to directly access the SSH service of the inference module (`10.10.99.98`). All model deployments should be performed through the storage card.
Before updating models or configurations, it is recommended to back up the `auto/` and `dev/` directories on the storage card to prevent configuration loss or the need to roll back to a previous version.
* Change all default passwords immediately after first login
* Use strong passwords (at least 12 characters, including uppercase and lowercase letters, numbers, and special characters)
* Update passwords regularly (recommended every 90 days)
* Do not use the same password across multiple systems
## User Management
### Adding Users
Use an administrator account to access the management interface
Click on the "User Management" option in the left navigation bar
Click the "Add User" button
Enter username, email, initial password, and permission level
Click the "Create" button to complete user addition
### Permission Settings
RM-01 supports granular permission control:
Full access permissions, including system settings
Can use all applications, manage their own data
Can use designated applications, restricted data access
Can only use specific demo applications
**Adjusting User Permissions**
1. Enter "User Management"
2. Select the user to modify
3. Click "Edit Permissions"
4. Set application access permissions and data access scope
5. Save changes
## Regular Maintenance
### Status Monitoring
Regular system status checks ensure smooth operation
After logging into the management interface, check the system status information on the dashboard:
* CPU usage
* Memory usage
* Storage space usage
* Data read/write speed
* Operating temperature
* Network connection status
### Update Management
Update models through the CFexpress Type-B storage card:
1. Obtain new model files from your service provider
2. Use a card reader to place the model files into the `auto/llm/` directory of the storage card
3. Ensure the file structure is correct (refer to the Model Management section)
4. Reinsert the storage card and restart the device
It is recommended to back up existing model files before updating
Contact your service provider for application update packages and specific update procedures.
Typically requires:
1. Log in to the application module (`10.10.99.99`)
2. Upload the update package
3. Complete the update process following the prompts
Contact your service provider for system update support.
System updates may require professional technical support. Do not attempt unauthorized system update operations on your own.
### Performance Optimization
If you notice device performance degradation, try the following steps
Go to "Application Management" and stop any applications you don't need
Go to "System Maintenance" and select the "Clear Cache" option
Select "Restart" in the management interface, or short press the power button and select "Restart"
### Troubleshooting
**Possible causes and solutions**:
* **Power issue**: Check power adapter connection, ensure using original power supply
* **Storage card issue**: Try removing and reinserting the storage card
* **System issue**: Hold the power button for 10 seconds to force shutdown, then restart
**Possible causes and solutions**:
* **Insufficient resources**: Close other running applications
* **Corrupted application**: Reinstall the application
* **Model issue**: Check if the required models are correctly loaded
**Possible causes and solutions**:
* **Network adapter issue**: Check the USB-C network adapter connection, confirm the host recognizes the AX88179A or RMinte RM-01 network interface
* **IP address configuration error**: Confirm the host IP address is `10.10.99.100` with subnet mask `255.255.255.0`
* **Firewall restrictions**: Confirm the enterprise firewall is not blocking the following addresses:
* `10.10.99.97` (Management Module)
* `10.10.99.98` (Inference Module)
* `10.10.99.99` (Application Module)
* **Network interface not enabled**: Confirm the RM-01 network interface is enabled in network settings
**Possible causes and solutions**:
* **Multiple applications running simultaneously**: Reduce the number of running applications
* **Large model load**: Switch to a smaller model
* **High temperature**: Ensure good device ventilation, avoid stacking
* **Insufficient storage space**: Clear unnecessary data
* **Automatic mode limitations**: If using automatic mode to load models, performance is limited to a maximum context of 8192 tokens. Contact your service provider to upgrade to manual mode for higher performance
**Possible causes and solutions**:
* **Incorrect file structure**: Ensure model files are placed directly in the `auto/llm/` directory, do not use subfolders
* **Incompatible file format**: Check if supported formats are used (`.safetensors`, `.bin`, `.pt`, `.awq`, etc.)
* **Missing required files**: Ensure complete model files are included, including `config.json`, `tokenizer.json`, and other configuration files
* **Storage card not properly inserted**: Reinsert the CFexpress Type-B storage card and restart the device
* **Storage card failure**: Try using another CFexpress Type-B storage card
## Getting Support
### Service Provider Contact Information
400-XXX-XXXX (Weekdays 9:00-18:00)
[support@rminte.com](mailto:support@rminte.com)
[https://support.rminte.com](https://support.rminte.com)
### Learning Resources
Comprehensive documentation and guides
Step-by-step visual learning materials
Connect with other developers and experts
Community discussions and Q\&A
Congratulations! You have completed the basic setup and usage overview of RM-01! Now you can:
* Use RobOS at `http://10.10.99.97` to monitor device status
* Access Open WebUI at `http://10.10.99.99` for model debugging
* Load your own open-source models to the `auto/llm/` directory via CFexpress Type-B storage card
* Use the inference service endpoint at `http://10.10.99.98:58000/v1/chat/completions` for AI inference
Start exploring the endless possibilities that AI technology brings to your enterprise! If you have any questions, please contact your service provider or visit our support center for assistance.
For in-depth development or optimizing RM-01 performance, please refer to the RM-01 Developer Guide or contact your service provider for additional technical support.
***
# Introduction
Source: https://docs.rminte.com/index
Welcome to the RM-01 AI Supercomputer Documentation - your comprehensive guide to deploying and managing enterprise AI solutions
## What is RM-01?
The RM-01 is a state-of-the-art portable AI supercomputer designed for private, on-premises deployment of large language models and AI applications. Built for enterprise environments, it delivers powerful AI capabilities while maintaining complete data sovereignty and security.
### Key Capabilities
Up to 2070 TFLOPS (FP4) computing power with 32/64/128GB GPU memory for efficient local AI inference
Dedicated 24GB RAM for AI applications and 8TB SSD for RAG data storage
Hardware-level asymmetrical encryption and secure key management providing fully private, on-premises operation with no data leaving your device
0 cost for deployment, maintenance and a flat learning curve making it the most cost-effective local AI solution for enterprises
## Core Features
### High-Performance Local Inferencing with Flexible Model Support
**Supported Model Sizes and Types:**
* 0.5B to 235B parameters (GPTQ Int4 quantization)
* Mainstream open-source models (DeepSeek, Qwen, Llama)
* Custom enterprise/industry-specific models
* Support for multi-modal models
**Inference Capabilities:**
* Concurrent multi-model support
* Multi-user support
* Hardware-accelerated inference with total throughput of over 2000 TPS (tokens per second)
**Deployment Guarantees:**
* Plug-and-play installation
* Real-time model switching
* Zero-maintenance design
* Enterprise integration ready
### One-stop Solution for Enterprise AI
* 24GB dedicated RAM for AI applications
* Managing and deploying AI applications with Docker
* Asymmetrical hardware-level encryption for safe application delivery
* Embedding model automatically loaded and ready for RAG applications
* Reranker model automatically loaded and ready for RAG applications
* Vector database automatically loaded and ready for RAG applications
* 8TB SSD for data storage
* One USB-C port for connection to any device
* RM-01 will be recognized as a network device
* Use and manage AI applications/models with assigned local IP addresses and APIs
* Could be used as a server with a USB-C to Ethernet adapter
### Enterprise-Grade Security & Privacy
All data processing occurs locally on your device. No information is transmitted to external servers or cloud services.
Built-in hardware-level encryption ensures your data remains secure at all times, even the device is lost or stolen.
Enterprise-grade user authentication and authorization systems protect against unauthorized access.
Meets industry standards for data protection and privacy regulations including GDPR, HIPAA, and SOC 2.
### Lightweight Deployment for Low Enterprise TCO
* One-click deployment within 300 seconds with 0 cost
* No dedicated IT staff required
* Not intrusive to existing IT infrastructure
* Extremely stable and reliable with 0 maintenance required
* Power off and on to recover from any unexpected issues
* No dedicated maintenance staff required
* Flat learning curve for enterprise users with no need to learn complex AI technologies
* Card swapping with one-touch deployment for any AI solutions
## What's Next?
Choose your path based on your role and requirements:
**For Enterprise Users**
Get started with your RM-01 device, including hardware setup, management interface access, and using AI applications.
**For Vendors & Service Providers**
Complete deployment guide for implementing RM-01 in enterprise environments, including model preparation and system integration.
**For Developers**
Technical guide for building applications on RM-01, covering system architecture, environment setup, and model integration.
***
# Technical Specifications
Source: https://docs.rminte.com/specifications
Technical specifications of RM-01 AI Supercomputer
### Performance at a Glance
**1,000 TOPS**
FP8 processing capability
**Up to 128GB**
Unified high-bandwidth memory
**TDP 100W**
98.6% Savings vs traditional AI servers
**0.5B - 235B**
Parameter range support
### Detailed Specifications
**Core Processing Capabilities**
* **AI Processing Power**: Up to 1,000 TOPS (FP8)
* **Precision Support**: FP8, FP16, FP32
* **Model Parameters**: 0.5B to 235B (GPTQ Int4 quantization)
* **Inference Speed**: Real-time with hardware acceleration
* **Concurrent Users**: Multiple simultaneous sessions supported
* **Model Switching**: Hot-swappable with zero downtime
* **Processing Latency**: Sub-second response times
**Memory Architecture**
* **GPU Memory Options**: 32GB / 64GB / 128GB configurations
* **Memory Type**: Unified high-bandwidth memory architecture
* **Memory Bandwidth**: Optimized for AI workload processing
* **Cache System**: Multi-level intelligent caching
* **Memory Management**: Intelligent resource allocation
**Performance Optimization**
* **Dynamic Batching**: Automatic request optimization
* **Model Quantization**: Advanced GPTQ Int4 support
* **Load Balancing**: Distributed processing capability
* **Throughput**: Enterprise-grade processing capacity
**Power Consumption Specifications**
Peak performance operation with full AI processing capabilities engaged
Normal operation power consumption during standard workloads
Estimated electricity cost for 24/7 continuous operation
**Energy Efficiency Comparison**
* **Traditional AI Server Power**: \~7,000W typical consumption
* **RM-01 Power Consumption**: 100W maximum
* **Energy Savings**: 98.6% reduction vs traditional infrastructure
* **Carbon Footprint**: Significantly reduced environmental impact
* **Cooling Requirements**: Independent cooling system, no additional cooling needed
**Cost Benefits**
* **Lower Initial Investment**: 80% savings compared to traditional setup
* **Reduced Operational Costs**: 98% decrease in ongoing expenses
* **3-Year TCO Savings**: 99% total cost of ownership reduction
**Storage Systems**
* **Primary Storage**: CFexpress Type B cards for model and application storage
* **System Storage**: 8TB SSD for data storage
* **Data Transfer**: High-speed model loading and switching
* **Storage Security**: Hardware-level encryption for all data
**Network & Connectivity**
* **Network Interface**: USB-C to Ethernet adapter included
* **Management Access**: Web-based management interface
* **API Support**: OpenAI Compatible API for system integration
* **Protocols**: Standard enterprise networking protocols
* **Security**: Enterprise-grade network security features
**Enterprise Integration**
* **Directory Services**: LDAP/Active Directory integration
* **Authentication**: Single Sign-On (SSO) support
* **Remote Access**: Secure VPN connectivity options
* **Monitoring**: SNMP and custom metrics collection
* **Compliance**: Enterprise security and audit requirements
**Physical Design**
Portable desktop design optimized for enterprise environments
Advanced independent cooling system for silent operation
Rugged construction for reliable operation
USB-C ports for power and network connectivity
**Environmental Specifications**
* **Operating Temperature**: -20°C to 60°C (-4°F to 140°F)
* **Storage Temperature**: -20°C to 60°C (-4°F to 140°F)
* **Operating Altitude**: Up to 5,000 meters above sea level
* **Certifications**: Meets enterprise environmental standards
**Deployment Considerations**
* **Installation**: Simple plug-and-play setup
* **Maintenance**: Zero-maintenance operation design
* **Portability**: Lightweight for easy relocation
* **Security**: Physical security features for enterprise use
### Performance Comparison
**Investment Savings**
* Initial hardware cost reduction: 80%
* Setup and deployment savings: Significantly reduced
* Training and implementation: Minimal requirements
**Operational Savings**
* Electricity costs: 98% reduction
* Maintenance costs: Near-zero maintenance required
* IT staff requirements: No dedicated IT staff required
**3-Year Total Cost of Ownership**
* Overall TCO savings: 99% compared to traditional AI servers
* ROI achievement: Typically within 1-3 months
* Ongoing costs: Predictable and minimal
**Data Privacy Advantages**
* Complete data sovereignty: All processing on-premises
* No data transmission: Zero cloud dependency
* Compliance ready: Meets strict data protection regulations
* Security control: Full enterprise control over AI operations
**Performance Benefits**
* Latency reduction: Local processing eliminates network delays
* Availability: No internet dependency for AI operations
* Customization: Full control over model selection and tuning
* Scalability: Predictable performance without usage limits
## Use Cases
### Enterprise Applications
Automated document analysis, summarization, and information extraction for enterprise workflows.
Intelligent chatbots and virtual assistants for customer support and internal help desk operations.
Automated content creation, technical writing, and marketing material generation.
Advanced analytics, pattern recognition, and insight generation from enterprise data.
### Industry Solutions
* Medical document analysis
* Clinical decision support
* Research data processing
* Compliance reporting
* Risk assessment modeling
* Fraud detection systems
* Regulatory compliance
* Market analysis tools
* Quality control automation
* Predictive maintenance
* Supply chain optimization
* Process documentation
* Contract analysis
* Legal research assistance
* Compliance monitoring
* Document review automation
## Why Choose RM-01?
Unlike cloud-based AI services, RM-01 keeps all your data on-premises, ensuring complete privacy and compliance with data protection regulations.
Eliminate recurring cloud costs and reduce total cost of ownership by up to 99% compared to traditional AI infrastructure.
Designed for enterprise environments with professional support, comprehensive documentation, and proven deployment methodologies.
From individual deployments to enterprise-wide implementations, RM-01 scales to meet your organization's needs.
## Next Steps
Begin with basic setup and start using your RM-01 immediately
Access technical support and additional resources
Need help getting started? Our technical support team is available to assist with deployment, development, and ongoing operations. Contact us at [support@rminte.com](mailto:support@rminte.com).
***