YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

AeroEmbed-1

AeroEmbed-1 project image

AeroEmbed-1 is AeroAI’s lightweight, aviation-oriented language-model project, described as a sub-1B-parameter model intended for relatively constrained devices.

Verify the published checkpoint, exact parameter count, architecture, and supported tool interface against the model card before presenting them as properties of a released build. This README documents the project’s training and evaluation design; it does not claim that every experiment below has been deployed in AeroEmbed-1.

Warning - AeroEmbed-1(92626)-GGUF is the first of the AeroEmbed family, with elevated errors, and a on average 9/12 pass rate on aviation queries. Its modified base, a open source 0.8B model is highly deprecated, and has elevated errors on basic Q+A questions, due to extended LoRa training.

Scope and safety

The exercise suite asks the model to read supplied values, perform comparisons and arithmetic itself at inference time, and give a concise, checkable explanation. It covers cylinder-head temperature (CHT), fuel, engine oil, manifold absolute pressure (MAP), and computational avionics examples.

The training-data generator and offline evaluator may use exact arithmetic to check labels. That is different from putting a calculator in the model’s inference path.

Safety boundary: These are synthetic exercises—not aircraft-specific operating instructions, approved engine limits, flight-planning outputs, or a primary engine, fuel, or navigation indication. Actual operating limits come from the applicable AFM/POH and installed-equipment documentation. Never substitute an experimental model’s generated result for a certified indication or required pilot checks. See the FAA Pilot’s Handbook of Aeronautical Knowledge, flight-manuals chapter, and Garmin fuel-computer documentation.

How it works

  1. Receive a task. The user supplies the exercise rules, units, thresholds, reserve, and requested output keys. Do not assume a universal CHT, oil, or MAP limit.
  2. Obtain measurements. A host application may supply a tool result, such as Get_Flight_Status, containing raw readings. In the current JSONL training examples, Tool result — Get_Flight_Status: {...} is text inside the user message. It is not evidence of an executed function call. A deployment with real tool calling must implement the actual call/result exchange and align its training serialization accordingly.
  3. Validate before calculating. Check status, required fields, units, numeric values, and any supplied engine cylinder count. A four-cylinder example requires four CHT readings; a six-cylinder example requires six. Zero fuel and zero groundspeed are valid exercise inputs. Zero fuel flow is invalid when the calculation requires division by flow. If a required input is unusable, return UNKNOWN, not a partially filled result.
  4. Compute in the model. Apply the specified strict inequalities, divisions, time conversions, subtraction, and distance multiplication. Do not round intermediate values. Groundspeed affects distance and estimated time en route (ETE), not fuel endurance in minutes for a given fuel amount and flow.
  5. Emit a verifiable answer. Write one Check: line exposing the comparison and arithmetic, followed by one Answer: line containing exactly the requested JSON object—or Answer: UNKNOWN for invalid inputs. A fluent Check: line is not proof that the math is correct; an independent offline evaluator must verify both lines.

For the avionics exercises, DME slant distance and horizontal distance are deliberately distinct: the FAA’s DME overview describes DME as measuring slant range. Oil and MAP thresholds are supplied per exercise, not inferred from generic aircraft rules.

Prompts and examples

System prompt for Version .9.2

Use the following prompt with Version .9.2-style examples. Keep it consistent between training and evaluation:

Use only the exercise rule and tool data. Validate status, required fields, units, and any stated engine cylinder count before computing. Treat zero fuel and zero groundspeed as valid, but zero fuel flow, missing readings, mismatched units, and stale data as invalid whenever required. Never infer or convert missing inputs. If any required input is invalid, write Check: with the reason, then Answer: UNKNOWN; do not compute a partial answer. Otherwise write one concise Check: line showing exact comparisons and calculations without rounding intermediate results, then one Answer: line containing only the requested JSON object. All limits and reserves are hypothetical exercise values, not aircraft operating limits.

This is an exercise-mode prompt. It does not authorize a live-flight advisory role. In a deployment with actual tools, do not pretend a tool ran when its result was merely written into a user message. If a task requires a live measurement and none is available, ask for it or indicate that the exercise cannot be completed.

CHT and fuel

Version .9.1/.9.2 core entries follow this structure. Replace the bracketed fields while keeping tool-result keys and units explicit:

Exercise CHT limit: [limit_f] °F. CHT is HIGH if any cylinder is strictly above the limit; otherwise FINE. Calculate fuel endurance in minutes as usable gallons divided by gallons per hour, times 60. Exercise reserve: [reserve_minutes] minutes. Minutes above reserve = max(0, endurance minus reserve). Calculate theoretical distance to empty in NM from current groundspeed and endurance hours. Use only current, complete data in the stated units. Return JSON with keys cht, endurance_minutes, minutes_above_reserve, and distance_to_empty_nm.

Tool result — Get_Flight_Status: {"status":"current","engine_type":"4-cylinder piston","cylinders_f":[437,441,450,448],"usable_fuel_gal":16,"burn_gph":8,"groundspeed_kt":100}

With limit_f = 450 and reserve_minutes = 45, the answer is:

{"cht":"FINE","endurance_minutes":120,"minutes_above_reserve":75,"distance_to_empty_nm":200}

Equality at the limit is not HIGH.

Additional question families

Each question should supply all necessary limits and formulas rather than ask the model to invent them. It also needs a corresponding status and correctly named measurement fields in its tool-result text. See the Version .9.2 JSONL for full prompt wording and exact JSON schemas.

Family Supplied exercise rule and example Expected result
Oil Pressure strictly below 50 psi is LOW; temperature strictly above 230 °F is HIGH. Compute pressure margin = current − 50, temperature margin = 230 − current, and pressure change = current − previous. Alert if either limit is violated. At 49 psi current pressure, 54 psi previous pressure, and 230 °F current temperature: oil_alert = true, pressure margin = −1 psi, temperature margin = 0 °F, and pressure change = −5 psi.
MAP At the supplied RPM, compare current MAP with the supplied upper limit in inHg using strict >. Compute margin = limit − current and change = current − previous. Do not infer power percentage from MAP alone. At 2400 RPM, with a 27 inHg limit, 27.5 inHg current, and 26.75 inHg previous: ABOVE, margin = −0.5 inHg, and change = +0.75 inHg.
Avionics math For hypothetical DME geometry, slant range = sqrt(horizontal_nm² + height_above_station_nm²). Separately, time to a fix = distance_to_fix_nm / groundspeed_kt × 60. Horizontal distance 1.2 NM and height 0.5 NM give 1.3 NM slant range. A separate fix 30 NM away at 120 kt gives 15 minutes ETE. Do not equate the two distances.

Dataset evolution

The stages below describe the design progression, with representative examples. They are not a claim that every earlier draft or checkpoint was independently archived or benchmarked under identical conditions.

Stage Representative training target Limitation discovered Revision
Plain-English questions “A cylinder reads 451 °F. Is that high?” The model may guess a universal limit or invent measurements; the task has no explicit input contract. Supply the exercise limit, strict inequality, and readings.
Tool-result questions Provide Get_Flight_Status text with four or six CHT readings and fuel fields. Parsing a tool result does not guarantee correct comparison, unit use, or arithmetic. The dataset’s tool-result text is not a real function-call trace. Train on explicit units, array length, and required fields; keep actual tool execution separate.
Model-computed reasoning Check: max = 451; 451 > 450 is true. ... followed by Answer: ... Plausible written steps can hide false comparisons, early rounding, or multiplication errors. Use short, numerically verifiable steps and independently checked labels.
Validation and checkpoints Check: validates status, units, and array completeness first; otherwise Answer: UNKNOWN. The model may complete an answer using a null reading or liters as gallons. Even a correct final answer may have a false Check: line. Train validation-before-math and targeted hard cases; score held-out checkpoints on both lines.

Concrete revisions

Preserve intermediate arithmetic. Do not silently round hours into minutes and let the rounded number affect a training label:

Earlier: Endurance = 13 ÷ 8 = 1.625 h, about 98 min; above a 35-minute reserve is 63 min.
Revised: Endurance = 13 ÷ 8 = 1.625 h; 1.625 × 60 = 97.5 min; max(0, 97.5 − 35) = 62.5 min.

Stop on incomplete data. A six-cylinder task cannot produce a full result from five readings:

Input: six-cylinder CHT readings [440, 451, null, 472, 456, 459].
Bad: ignore null, calculate from five readings, return a full JSON result.

Revised Check: One of the six CHT readings is null; the required data is incomplete. Stop; no partial answer.
Revised Answer: UNKNOWN

Reject mismatched units. Treat usable_fuel_liters as a unit mismatch, not as usable_fuel_gal. Do not divide a value labeled liters by a gallons-per-hour flow merely because both fields contain numbers. The exercise prompt disallows guessing or converting missing required fields.

Dataset lineage

Revision Contents
Initial worked example Explicit Check: arithmetic and structured Answer: JSON for four- and six-cylinder CHT plus fuel.
60-entry dataset Straightforward arithmetic, equality boundaries, and invalid-input examples.
Version .9 480 entries: 360 valid answers and 120 UNKNOWN answers.
Version .9.1 Kept 480 entries while clarifying validation rules, exact arithmetic, and some expanded multiplication checks.
Version .9.2 A combined replacement file—not a second file to append to .9.1—with 864 entries: 480 CHT/fuel-family entries and 384 added examples across CHT trends, fuel-to-fix, oil, MAP, avionics, and integrated multi-system tasks. Its outputs and calculations were checked when generated. Use the system prompt embedded in the file.

Evaluation and release

Checkpoint findings

A short checkpoint run in this development conversation returned 8/12 correct final answers on the C01–C12 set. That alone does not establish improvement over an earlier reported 5/10 result: the question sets differ, and the log does not establish that the same base model, decoding settings, and training status were used. Do not attribute the difference solely to dataset format.

The four incorrect final answers identify useful training targets:

Checkpoint Failure Correct treatment
C07 Converted 1.375 hours to 81 minutes. 1.375 × 60 = 82.5 minutes; after a 30-minute reserve, 52.5 minutes remain.
C08 Calculated 3.125 × 104 as 326.325 NM. 3.125 × 104 = 325 NM at the stipulated constant groundspeed.
C10 Skipped a null among six CHT readings and produced a full answer. Return UNKNOWN because a required reading is absent.
C12 Divided a number labeled liters by a flow labeled gallons/hour. Return UNKNOWN under the no-conversion exercise contract.

For every checkpoint, score:

  1. Whether Check: correctly validates the inputs.
  2. Every intermediate comparison and numerical expression.
  3. The final JSON fields and units.
  4. Whether invalid inputs produce only UNKNOWN.

Keep C01–C12 out of training, including lightly reworded copies or examples with the same distinctive measurement arrays. Build larger, new held-out sets: twelve examples produce a noisy estimate. Compare old and new model versions on the same untouched test items and decoding settings. Offline answer checking is allowed; the inference-time model is still expected to do the math itself.

Training and release notes

The JSONL structure is:

{"messages":[{"role":"system","content":"..."},{"role":"user","content":"..."},{"role":"assistant","content":"Check: ...\nAnswer: ..."}]}

The assistant target contains a visible, concise Check: line followed by Answer: JSON or UNKNOWN. The format does not require publishing an unbounded hidden chain of thought.

  • Keep the output contract consistent within an experiment. If a real tool-call API is added, document its schema and distinguish the assistant’s tool request, the tool’s response, and the assistant’s final answer. The current user-message text convention is not a substitute for tool-call supervision.
  • When training successive versions, log the base checkpoint, dataset revision, train/validation split, training method, learning rate, number of steps, decoding settings, and test-set hash. Otherwise, score changes cannot be attributed reliably.
  • For the small exercise dataset, compare LoRA before committing to full fine-tuning. A larger set and a held-out test are more informative than training loss alone. The training method does not add a calculator to inference.
  • Treat a successful model-only math benchmark as research progress, not evidence of approval for aircraft installation or use as the sole source of a flight-critical indication. Aircraft-specific limits and software assurance are separate engineering requirements.

Base model benchmark scores:

Benchmark Non-thinking Thinking
MMLU-Pro 29.7 42.3
MMLU-Redux 48.5 59.5
C-Eval 46.4 50.5
SuperGPQA 16.9 21.3
IFEval 52.1 44.0
MMMLU 34.1 44.3
GPQA — 11.9
IFBench — 21.0
LongBench v2 — 26.1
BFCL-V4 (tool use) — 25.3
TAU2-Bench (agent tasks) — 11.6

Warning AeroEmbed-1(92626)-GGUF is the first of the AeroEmbed family, with elevated errors, and a on average 9/12 pass rate on aviation queries. Its modified base, a open source 0.8B model is highly deprecated, and has elevated errors on basic Q+A questions, due to extended LoRa training.

Downloads last month
60
GGUF
Model size
0.8B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support