SuperFast Tiny Home Robotic

An offline English and Hinglish command interpreter for home devices and registered industrial/robotics devices. It converts one instruction into a catalog-validated JSON intent. The package includes a trained 87,181-byte action model, a 1,002,414-device SQLite catalog, a Python API, an HTTP service, a browser playground, Docker files, and reproducible training and evaluation data.

Publisher: sraivante. Release: v1.0.0, 2026-09-26. The runtime was previously packaged locally as edge-command-docker-v3-ui.

stop the robot arm 2 at factory 1 line 1 station 1
  -> {"activity":"robotics","subject":"factory_1_line_1_station_1_robot_arm_2","action":"STOP"}

The system resolves device names through explicit identifiers, aliases and SQLite FTS5 search. It prioritizes action patterns and capability checks, with a stored count-based unigram/bigram classifier for the fallback scoring path. The evaluation below measures the complete hybrid interpreter, not the learned classifier in isolation. No external base model or neural checkpoint is used. This is a custom Python package; load it with Interpreter.

Files and requirements

Component Size / scope
action_model.json.gz 87,181 bytes compressed; 180,309 training rows; 6,892 retained features; 12 action classes
catalog.sqlite3 167,702,528 bytes; 1,001,410 base devices plus 1,004 legacy home devices
legacy_subjects.json.gz 3,863 bytes; canonical legacy identifiers
Runtime Python 3 with SQLite FTS5; standard library only; no GPU or network required after download
edge_command.py, build.py Inference, features and original model/catalog construction
server.py, index.html, compose.yaml, Dockerfile HTTP API, browser UI and deployment
evaluation.json, release_evaluation.json Original diagnostic and release-time rerun
training_snapshot.json, release_manifest.json, SHA256SUMS Provenance, pinned dataset reference and file integrity
config.json Release metadata; also the file Hugging Face uses to count downloads (added in v1.0.1, weights unchanged)
edge-command-docker-v3-ui.tar.xz, ORIGINAL_README.md Original supplied package and its documentation

The tiny figure describes the action model. The catalog is a separate 167.7 MB on-disk dependency. SuperFast is the release name, not a measured comparison with other models. Raspberry Pi latency and process memory are now recorded below. Field deployment and end-to-end hardware performance remain unmeasured.

Raspberry Pi 5 edge-device performance

Measured 2026-09-26 on a Raspberry Pi 5 Model B, 16 GB RAM, Debian 13.2/aarch64, Python 3.13.5. Measurement update: edge-rpi5-20260926.1; weights and training data are unchanged.

Threads Median ms p95 ms p99 ms Serial requests/s Process RSS MiB CPU °C min–mean–max Timed correct
1 1.28 2.58 2.60 763.73 27.19 59.5–61.6–63.4 500/500

Batch 1; 20 warmups and 500 timed calls per row, cycling over 10 frozen model-specific inputs. Each timed reading has a matching CPU-temperature sample, captured immediately afterward outside the latency timer; the table shows min–mean–max temperature. RSS includes Python/runtime overhead. Timings include local text-to-output inference and exclude loading, SSH/HTTP, speech recognition and actuation. Throughput is serial reciprocal mean. This catalog interpreter is single-threaded. These synthetic workloads do not establish cross-model superiority or robot-control reliability.

On-device quality check: 1274/1346 on existing template diagnostic and exact-ID/rejection probes. See the report for per-category results and previously observed limitations.

Full Pi methodology, temperature plot and reproduction · Every latency/temperature reading · Raw measurements · Versioned evaluation data.

Download and run

pip install huggingface_hub
python -c "from huggingface_hub import snapshot_download; snapshot_download('sraivante/superfast-tiny-home-robotic', revision='v1.0.1', local_dir='superfast-tiny-home-robotic')"
cd superfast-tiny-home-robotic
python edge_command.py "stop the robot arm 2 at factory 1 line 1 station 1"

After downloading, inference works offline. From the repository directory:

from edge_command import Interpreter

engine = Interpreter()
try:
    print(engine.parse("turn on kitchen light"))
    # {"activity": "light", "subject": "kitchen_light", "action": "ON"}
finally:
    engine.close()

Start the UI and HTTP API:

docker compose up --build -d
# Open http://127.0.0.1:8080/
curl -s http://127.0.0.1:8080/healthz
curl -s -X POST http://127.0.0.1:8080/parse \
  -H 'Content-Type: application/json' \
  -d '{"instruction":"turn on kitchen light"}'

POST /parse accepts {"instruction":"..."} and returns a parsed intent with HTTP 200 or a JSON error with HTTP 422. GET /health and GET /healthz check readiness. Compose binds the host port to loopback. Without Docker, python server.py starts the same service on port 8080 and binds to all interfaces; use it on a trusted local network. The service has no authentication.

Tasks, languages and output

Supported input is English and Romanized Hindi/Hinglish. Devanagari, speech recognition, conversation context and translation are not implemented.

{"activity":"robotics","subject":"factory_1_line_1_station_1_gripper","action":"START","requires_confirmation":true}

Successful responses contain activity, a registered subject, and a catalog-supported action. The action set is ON, OFF, LOW, MEDIUM, HIGH, OPEN, CLOSE, LOCK, UNLOCK, START, STOP, STATUS. Sensitive activity categories, except STOP, and every UNLOCK request add requires_confirmation: true. Failures return an error string and may include candidate identifiers or a subject.

The interpreter produces intents only. It does not move hardware or obtain live device state. STATUS requests a future status query; it is not a measurement. A downstream controller must validate authorization, device capabilities, state and interlocks. A parsed STOP is not an emergency stop.

Training data and provenance

The matching data is published at SuperFast Tiny Home Robotic Data. The immutable dataset commit is 5d997497d2cbd01f7104babd576498a9fe1f0975.

The original build.py thins the 2,141,541-row command source in file order with random.Random(42) and selection probability 180000 / 2141541. It selects 180,309 rows, counts unique word/number unigrams and adjacent bigrams per record and action, and retains features occurring at least three times subject to its length filter. Training is one streaming count-aggregation pass, with no optimizer, gradient steps, epochs beyond that pass, or fine-tuning stage. Runtime scoring uses count-plus-one smoothing. Source identifiers are not explicitly removed from features. Original build hardware was not recorded.

For this release, replaying that procedure reproduced every stored class total, feature count, feature total, vocabulary size and training-row count exactly. All legacy identifiers and non-timing evaluation results also match. The source, selected rows, sampled diagnostic rows, catalog probes and rejection cases are included with hashes and source-row numbers.

No contemporaneous raw-source hash was stored with the original run. These matches provide strong reproducibility evidence, while leaving the original source-byte identity independently unproven. The missing original devices-1M.jsonl.xz is represented by a lossless logical export of its stored device fields from catalog.sqlite3; original formatting, discarded fields and the catalog generator are unavailable. These distinctions are recorded in training_snapshot.json.

Evaluation

Diagnostic Exact / total Result
Sampled home-command triples 928 / 1,000 92.8%
Registered catalog action/ID probes 340 / 340 100%
Expected rejection cases 6 / 6 100%

The release-time rerun reproduced the original results and saved all predictions in the dataset repository. The template sample uses reservoir sampling with seed 198 from the same source used for training; 82 of 1,000 rows are also selected training rows. There is no independent held-out evaluation. Exact match checks the three target fields and permits an extra confirmation flag. Catalog probes use exact registered IDs and exercise capability coverage. The six rejection cases provide narrow regression coverage. Neither establishes generalization to real speech or safe hardware operation.

The template diagnostic includes 15 incorrect triples, 29 ambiguous-device rejections, 23 unknown/unspecified-device rejections and 5 unsupported-action rejections. Sample failures and the full predictions are published.

Reproduce

Download the matching dataset beside this repository:

python -c "from huggingface_hub import snapshot_download; snapshot_download('sraivante/superfast-tiny-home-robotic-data', revision='5d997497d2cbd01f7104babd576498a9fe1f0975', local_dir='../superfast-tiny-home-robotic-data')"
python build.py ../superfast-tiny-home-robotic-data/source/devices-1M-recovered.jsonl.xz ../superfast-tiny-home-robotic-data/source/home-commands-v2-part1.jsonl.xz --out rebuilt
python evaluate.py ../superfast-tiny-home-robotic-data/source/home-commands-v2-part1.jsonl.xz --output evaluation-rerun.json
python edge_command.py --path rebuilt "turn on kitchen light"

evaluate.py evaluates the packaged model in its own directory. The last command separately exercises the rebuilt directory. Compare decompressed model JSON and logical catalog rows when checking a rebuild: gzip timestamps, dictionary ordering and SQLite versions can change physical file hashes.

Limitations

  • Synthetic templates dominate the command data. Novel wording, ASR errors, ambiguous aliases and conflicting instructions can fail or return wrong intents.
  • Negation and multiple-command rejection use limited patterns. They are not comprehensive language understanding or a security boundary.
  • Supports a single catalog-constrained command. Conditional rules, motion trajectories, destinations, numeric speed values and manipulation planning are not represented by this package.
  • The catalog contains predefined identifiers and capabilities. New devices require an explicitly configured catalog; it is not live device discovery.
  • No hardware operation, permissions system, telemetry integration or real-time control loop is provided. No comparative speed or field-reliability claim is made.

License and attribution

Apache License 2.0 applies to the user's original model, code, dataset material and selection/arrangement. Copyright (c) 2026 sraivante. Third-party ownership, licenses and runtime dependencies remain separate; the copyright statement does not claim them. No third-party model weights are included. See LICENSE and NOTICE. Dataset generator provenance and remaining historical uncertainty are documented in the linked dataset card.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train sraivante/superfast-tiny-home-robotic

Evaluation results

  • Exact activity, subject and action match (hybrid interpreter) on Template diagnostic (1,000 rows; overlaps training)
    test set self-reported
    92.800