# Iteration 2: Unified FLOSS pipeline and schema
**Status:** ready to implement
**Baseline:** PR [#1337](https://github.com/mandiant/flare-floss/pull/1337) (Iteration 1 modular split)
**Form:** one PR, ordered commits (not a multi-PR stack)
**Branch context:** `quantumstrand` / first-class layout+tags modules already landed
## Goal
One analysis path, one `ResultDocument`, one default human and JSON output that includes deobfuscation and layout/tags. End the dual product (`floss` vs `floss quantum`).
When Iteration 2 is done:
```text
$ floss sample.exe # layout+tags static (and language) when supported
# stack / tight / decoded at the end, unenriched
$ floss sample.exe -j # one JSON schema; enrichment on static/language only
# DB path: floss/tags/data (LFS); no floss/qs/db
# no required: floss quantum
```
## Locked product decisions
| Topic | Decision |
|-------|----------|
| Compatibility | Major release; no legacy JSON/API compatibility |
| Default UX | Layout-aware static output when layout succeeds (quantum-style) |
| Feature flag | No permanent `--quantum`; optional `--no layout` / `--no tags` only |
| Naming | Plain fields: `tags`, `section`, `structure`, `layout` |
| Binary | One `floss` binary; collapse `quantum` entry |
| Tag DBs | Move `floss/qs/db/data` → `floss/tags/data` in this iteration |
| GUI | Out of scope (`qs-viewer` stays unmigrated) |
| Enrichment scope | **Static + language only**; stack / tight / decoded unenriched |
| Static extract | Quantum extract when layout succeeds; classic fallback when it fails |
| PR shape | Single PR, separate commits |
## Current state (after #1337)
Done:
- First-class `floss/layout/`, `floss/tags/`, `floss/ranges.py`, `floss/cli.py`
- QS monolith removed; `build-qs.yml` / `qs.spec` dropped
- Core deps always installed
- OSS DB text diffs in PR bodies
- PyInstaller packs tag DBs (still under legacy path)
- `pipeline.py` is a stub pointing at Iteration 2
Still dual:
| Classic FLOSS | Quantum path |
|---------------|--------------|
| `floss sample` → `main.py` analysis body | `floss quantum` / `quantum` → `quantum.py` |
| `floss.results.ResultDocument` | `floss.document.ResultDocument` |
| Deobfuscation, no layout/tags | Layout + tags, static only |
| Tag DBs at `floss/qs/db/data` | Same path via `data_root()` |
## Architecture
```mermaid
flowchart TD
sample[Sample path]
main[floss.main
parse args, logging, exit codes]
pipe[floss.pipeline.analyze]
static[Static / language extract]
layout{Layout supported
and enabled?}
qextract[Quantum extract inside layout
tag, mark structures, FP filter]
classic[Classic get_static_strings]
enrich[Project tags / section / structure
onto static and language strings]
viv[Vivisect stack / tight / decoded
if enabled]
doc[Single ResultDocument]
text[Text render]
json[JSON render]
sample --> main --> pipe
pipe --> static --> layout
layout -->|yes| qextract --> enrich
layout -->|no / fail| classic --> enrich
enrich --> viv --> doc
doc --> text
doc --> json
```
Target package layout:
```text
floss/
main.py # thin CLI entry
cli.py # argparse
pipeline.py # full analysis orchestration
results.py # unified schema (layout + tags + deobfuscated strings)
ranges.py
layout/ # PE / ELF / Mach-O
tags/ # taggers + data/ (LFS)
render/
default.py # hybrid human output
json.py
layout_text.py # section tree for static when layout present
```
Deleted or gutted by end of the PR:
- Dual `floss/document.py` schema (folded into `results.py`)
- Standalone product role of `floss/quantum.py` (alias then remove)
- Residual `floss/qs/` after DB move
## Schema
Evolve `floss.results` (not promote `document.py` as the long-term root).
```text
ResultDocument
metadata path, version, imagebase, min_length, language*, runtime*, …
analysis enable_* flags, function stats
optional enable_layout / enable_tags
layout optional serializable tree | null
strings
static_strings
language_strings
language_strings_missed
stack_strings
tight_strings
decoded_strings
```
Per-string enrichment fields (plain names):
| Field | Type | Who gets values in Iteration 2 |
|-------|------|--------------------------------|
| `tags` | `list[str]` | static, language (+ missed) |
| `section` | `str` | static, language when layout known |
| `structure` | `str` | static, language when in a known structure |
Stack / tight / decoded keep empty defaults (`tags=[]`, `section=""`, `structure=""`).
Canonical string lists live under `strings.*`. The layout tree is for section-aware render and structure context, not a second divergent copy of string text after projection.
## Static extraction policy
```mermaid
flowchart LR
A[Start static analysis] --> B{compute_layout OK?}
B -->|yes| C[extract_layout_strings]
C --> D[tag_strings]
D --> E[mark_structures]
E --> F[remove_false_positive_lib_strings]
F --> G[Replace strings.static_strings
from layout]
B -->|no| H[get_static_strings fallback]
H --> I[No layout tree / empty enrichment]
G --> J[Enrich language strings by offset]
I --> J
```
Rules:
1. **Layout succeeds** — reuse quantum extract inside layout; apply tags, structures, hide junk / FP-lib rules; that set becomes `strings.static_strings`.
2. **Layout fails, unsupported format, or `--no layout`** — classic `get_static_strings`; no layout tree (or null).
3. **Codify** against PMA and existing `tests/test_qs_pma0101.py` expectations so static output does not regress vs quantum.
Language strings still come from Go/Rust extractors; enrich via layout offset mapping (tags / section / structure), not via re-extract through layout gaps unless already required for parity.
## Tagging scope
**No formal content vs layout tagger split** in this iteration. That was not migration scope; it only mattered for recovered-string enrichment, which we are not doing yet.
Keep existing:
- `load_databases()`
- `layout.tag_strings(...)`
- structure marking and FP-lib filter
Optional thin helpers only to avoid duplication (for example offset → section name). Do not invent a new tagger taxonomy.
Later iteration may add content-only tagging for stack / tight / decoded.
## Default text output
### Layout present (and static enabled)
Match quantum: **no classic meta table**.
```text
── .text ─────────────────────────┐
… tagged static strings …
── .rdata ────────────────────────┤
…
└ … layout tree …
[language section if applicable]
FLOSS STACK STRINGS (…) # if enabled / present
…
FLOSS TIGHT STRINGS (…)
…
FLOSS DECODED STRINGS (…)
```
Order:
1. Layout-aware static (Rich section tree via `layout_text`)
2. Language (if any; keep simple unless tests demand more)
3. Stack → tight → decoded at the end (classic headings and list style, unenriched)
Optional one-line `FLARE FLOSS RESULTS` banner is unnecessary; quantum never had it. Recovered-string headings alone mark those blocks.
### Layout absent
Classic meta table + classic static listing, then language and recovered sections as today.
### JSON
Always one schema. Metadata remains in the document either way. Enrichment fields populated only for static / language when layout+tags ran.
## Commit plan (single PR)
Each commit should be reviewable and ideally green on its own.
### Commit 1 — Schema (A)
- Extend string types in `results.py` with `tags`, `section`, `structure` (defaults)
- Add optional serializable `layout` tree on `ResultDocument`
- JSON round-trip tests only; do not duplicate existing layout/DB suites
- Behavior unchanged for classic CLI
### Commit 2 — Enrichment helpers (B′)
- Map layout results onto static / language string fields
- No tagger split
- No stack / tight / decoded enrichment
- Helpers such as offset → section / structure if useful
### Commit 3 — Pipeline extract (C)
- Move classic analysis body from `main.py` into `pipeline.analyze`
- `main` becomes parse → analyze → render → print
- Behavior-neutral; quantum still separate until later commits
### Commit 4 — Wire layout into pipeline (D)
- When layout OK: quantum extract → replace `static_strings`; tag / structure / FP rules
- Fallback classic statics when layout fails
- Enrich language strings by offset
- Analysis flags: layout/tags on by default for PE/ELF/Mach-O; shellcode/unknown skip cleanly
- Parity tests: PMA / `test_qs_pma0101.py` (and other quantum static expectations that would regress)
### Commit 5 — Unified default render (E)
- Layout present: quantum-style static tree, **no meta**
- Language then stack / tight / decoded at the end (classic blocks)
- Layout absent: classic meta + classic statics
- Quiet / no-color paths remain usable
### Commit 6 — DB path move
- `git mv floss/qs/db/data` → `floss/tags/data` (preserve LFS)
- Update:
- `.gitattributes` LFS patterns
- `floss.tags.data_root()`
- `.github/pyinstaller/floss.spec` datas
- `MANIFEST.in` if needed
- `scripts/tags/*`
- `.github/workflows/build-oss-db.yml` and related paths
- tests and docs that hardcode `qs/db`
- Delete residual empty `floss/qs/` when nothing remains
- Do not re-add LFS blobs as normal files
### Commit 7 — Collapse quantum (F)
- `floss quantum` → deprecation alias to the same pipeline, or remove if acceptable
- Remove `quantum` console script from `pyproject.toml` (or deprecate for one release)
- Delete or gut `floss/quantum.py` and interim `floss/document.py`
- CLI help and README: unified tool; default layout/tags; `--no layout` / `--no tags`
- Point tests at pipeline / results (renaming `test_qs_*` can wait for Iteration 3)
```mermaid
gitGraph
commit id: "1 schema"
commit id: "2 enrich helpers"
commit id: "3 pipeline extract"
commit id: "4 wire layout"
commit id: "5 default render"
commit id: "6 tags/data move"
commit id: "7 collapse quantum"
```
Commit 6 may land after commit 4 once `data_root()` is the single path choke point; keep it before or with F so the PR does not ship dual DB locations.
## CLI surface
| Invocation | Behavior |
|------------|----------|
| `floss sample` | Default: layout+tags when supported; recovered strings if enabled |
| `floss sample -j` | Single enriched schema |
| `floss --no layout …` | Classic static path; no tree |
| `floss --no tags …` | Layout optional; empty tags |
| `floss --only stack …` | Unchanged string-type selection |
| `floss quantum …` | Alias/deprecated during transition, then gone |
## Testing
Keep green (adapt imports/paths as needed):
- `tests/test_qs_*.py` (layout PE/ELF/Mach-O, code ranges, PMA0101, unit)
- Tag DB tests (`test_gp_db`, `test_winapi_db`, `test_qs_oss_db`, …)
- `tests/test_render.py`, `tests/test_load.py`
- Classic extraction tests
Add or extend only what is necessary:
1. Schema JSON round-trip with enrichment fields (commit 1)
2. Pipeline / static parity with PMA and `test_qs_pma0101.py` (commit 4)
3. Render smoke: layout present → no meta, static tree; recovered sections at end (commit 5)
4. Fallback: non-PE / layout fail → classic path, no crash
5. Flags: `--no tags`, `--no layout`
6. After DB move: paths resolve under `floss/tags/data`
Gates:
```text
pytest
pytest -k "layout or tags or qs or pipeline or main or render"
# workflows and pre-commit per AGENTS.md
```
## Explicitly out of Iteration 2
- Content tagging of stack / tight / decoded
- Formal tagger architecture split
- Broad new corpora (disk vs dump, fat Mach-O) beyond existing tests
- Deleting `qs-viewer` / `web-release.yml` (still out of scope)
- OSS CI matrix cost tuning (libmysql, etc.)
- Major version bump packaging and full release notes (minimal README/CLI blurb in commit 7 is enough)
## Iteration 3 handoff
After this PR:
1. Rename remaining `test_qs_*` / experiment naming polish
2. Remove dead alias code if quantum was only deprecated
3. Major version / changelog of JSON break
4. CI matrix and broader edge-case corpus
5. Any residual docs cleanup
## Success checklist
- [ ] One `ResultDocument` schema; interim dual document gone
- [ ] `pipeline.analyze` owns orchestration; `main` thin
- [ ] Layout success → quantum static behavior (junk/FP rules); fallback classic
- [ ] Enrichment on static and language only
- [ ] Default text: no meta when layout present; recovered strings at end
- [ ] JSON uses plain `tags` / `section` / `structure` / `layout`
- [ ] DBs at `floss/tags/data` with LFS, scripts, CI, and `floss.spec` updated
- [ ] No required `floss quantum` for normal use
- [ ] Existing tests adapted and green; round-trip + PMA parity covered
- [ ] Single PR with the commit series above