Technical proposal

NEURO-SYMBOLIC CODING
ARCHITECTURE

A Verification-First System for Full-Repository Software Engineering

Download .txt

~17 kB · plain text · 11 sections


1. EXECUTIVE SUMMARY AND PARADIGM SHIFT

This proposal describes a repository-aware coding system that combines a compact neural proposer with deterministic program analysis and execution. The central shift is from treating code generation as a contest to scale raw model parameters to treating it as a controlled search over candidate changes, where every candidate is checked against the actual repository and its behavior.

Why not a pure non-LLM math formula on two NVIDIA B300 GPUs? A formula or fixed symbolic procedure can be excellent within a narrowly specified domain, but full-repository coding requires interpreting ambiguous requests, discovering intent across many files, proposing unfamiliar implementations, and adapting to diverse frameworks and conventions. These tasks have no single compact closed-form solution. GPUs can accelerate learned inference and parallel candidate evaluation, but more compute does not turn a hand-authored formula into a general code-understanding system. The practical role for deterministic methods is not to replace proposal and interpretation; it is to constrain and verify them.

Why verification loops can outperform a frontier model used alone, including a hypothetical "Opus 5.5" at full-repository coding: a standalone model may produce plausible code while inventing symbols, violating local APIs, overlooking tests, or misunderstanding cross-file effects. A verifier has direct access to the checked-out repository, compiler, language server, tests, and runtime. By generating a patch, checking it, and feeding concrete failures back into the next attempt, the system turns errors into actionable evidence. This is an architectural hypothesis, not a claim of universal superiority: performance depends on repository indexing, test quality, task scope, and latency budget. The model supplies flexible proposals; deterministic tools supply grounded constraints.

The system's objective is not merely to maximize tokens per second. It is to maximize the rate of accepted, correct repository changes under a bounded compute and wall-clock budget. A candidate that is rejected quickly by compilation or tests can be cheaper and safer than a fluent but unverified answer.

2. DUAL-STATE SYSTEM ARCHITECTURE

The system maintains two complementary states:

A. Semantic/repository state: the current git tree, parsed syntax trees, symbol and reference indexes, call graph, inferred types, diagnostics, tests, and execution artifacts. This state is grounded in files and tools.

B. Proposal/search state: the task specification, invariant constraints, candidate patches, feedback history, and search scores. This state is used to decide what to try next. It must never override repository facts supplied by the deterministic layer.

2.1 Generator/proposer layer

Use a compact or grafted model in the 14B-27B parameter range as a patch proposer. Its role is deliberately bounded: it proposes small, reviewable unified diffs that satisfy an explicit task and a set of repository-derived invariants. It does not declare a patch correct, invent repository facts, or directly mutate the canonical working tree.

The proposer receives a curated context package: the user request; relevant files and surrounding code; definitions and usages returned by the language server; applicable interfaces and types; nearby tests; repository conventions; and the current candidate's diagnostics. Output should be machine-parseable and include a patch plus a concise rationale and any assumptions. Reject malformed diffs, out-of-scope paths, oversized changes, and edits that violate policy before execution.

Use retrieval and context selection from the local index rather than dumping the entire repository into the prompt. Prefer narrow patches, explicit invariants, and staged edits. Keep the original git tree immutable during candidate evaluation; each candidate applies in an isolated worktree or container.

2.2 Deterministic verifier layer: headless LSP daemon

Run a headless Language Server Protocol daemon for each supported language or workspace. The verifier should expose machine-readable operations for:

This layer reduces repository-level hallucination by answering "does this symbol exist here, what does it mean, and what depends on it?" from the checked-out code. It cannot prove all runtime behavior, and dynamic languages or reflection may limit completeness. Therefore, report uncertainty and fall back to tests and execution rather than treating an incomplete index as proof.

Maintain an incremental index keyed to git tree/content hashes. After a candidate patch, update only affected documents and dependent indexes where possible. Keep the LSP daemon warm to avoid repeated startup and indexing costs. Record the language-server version, workspace configuration, and diagnostics with each evaluation for reproducibility.

2.3 Execution sandbox

Evaluate every candidate in an ephemeral, isolated container or equivalent sandbox with bounded CPU, memory, disk, network access, and wall-clock time. The harness should:

  1. 1.Apply the patch to a disposable worktree.
  2. 2.Run formatting and lint checks configured by the repository.
  3. 3.Compile or type-check the affected package and, when affordable, the full project.
  4. 4.Run targeted tests first, then relevant existing suites, and optionally broader suites within the remaining budget.
  5. 5.Capture exit codes, compiler stdout/stderr, stack traces, test output, resource use, and timeout status.
  6. 6.Check task-specific runtime invariants and preserve artifacts for debugging.

Use repository-native commands rather than assuming one stack. Examples include pytest for Python, vitest for JavaScript/TypeScript, and cargo test for Rust. Discover commands from project configuration and existing CI where possible. Never run untrusted candidate code on the host or with ambient secrets. Network should be disabled by default, with narrowly scoped exceptions where a task requires it.

3. MULTIMODAL VISUAL-DIFF CLOSED LOOP FOR FRONTEND REPLICATION

For UI replication, combine browser rendering and image comparison with the code-verification loop.

3.1 Target ingest and rendering

Ingest the target screenshot or reference page and record its dimensions, device scale factor, fonts/assets assumptions, and viewport. Render the candidate page in headless Playwright at standardized desktop 1080p and mobile viewports. Normalize browser version, viewport, device scale, color scheme, and font availability so repeated comparisons are meaningful. Save screenshots and browser console/network errors as evaluation artifacts.

3.2 Perceptual comparison

Align reference and candidate images before scoring. Compute a combination of pixel/perceptual loss, SSIM (structural similarity), and CLIP-based visual similarity. No single score is sufficient: pixel loss is sensitive to antialiasing and minor shifts; SSIM captures local structural similarity; CLIP captures higher-level semantic resemblance but may miss precise geometry. Use weighted, calibrated metrics and retain per-region scores rather than only a global score.

A spatial error engine should identify regions with meaningful residuals and emit bounding boxes with coordinates, severity, and likely category (layout, typography, color, image, spacing, or missing element). Map image coordinates back to viewport coordinates and, where browser instrumentation allows, associate regions with DOM elements and computed styles.

3.3 Feedback loop

Feed the proposer a compact visual context package: the target and current screenshots, metric deltas, bounding boxes, relevant DOM/element metadata, and the CSS/HTML/asset definitions controlling those regions. Ask for a minimal adjustment, then rerender and recompute. Stop when the calibrated perceptual delta falls below the agreed threshold, the budget is exhausted, or progress stalls. Apply per-viewport acceptance thresholds so a desktop improvement does not silently break mobile.

Every visual candidate still passes ordinary parsing, type/lint checks, and relevant tests. Guard against overfitting to one screenshot with multiple viewports, representative routes, and checks for responsive behavior. Record the threshold and metric weights as configuration, not hidden assumptions.

4. TEST-TIME SEARCH AND VERIFICATION: MCTS / TREE OF THOUGHTS OVER GIT STATES

Represent search as a graph whose nodes are candidate git tree states and whose edges are proposed unified diffs. A node is identified by its parent state and patch content (or resulting tree hash), and stores diagnostics, test outcomes, visual scores where relevant, resource cost, and provenance. The canonical branch remains untouched until a candidate is accepted.

At each expansion, use the proposer to produce a small set of meaningfully different candidate patches, conditioned on the task and evidence collected so far. Search can use Monte Carlo Tree Search, best-first search, or a Tree-of-Thoughts-style controller; the specific policy should be selected empirically. Use explicit branching and depth limits to prevent unbounded generation.

Deterministic pruning is the primary efficiency lever. Immediately truncate a candidate path on patch-application failure, syntax/compilation failure, mandatory linter failure, or failure of required baseline tests, unless the task explicitly changes the failing behavior and the evaluator can distinguish expected from unrelated failures. Run cheap checks before expensive suites. Cache results by tree hash, command, environment, and test configuration so identical states are not reevaluated.

Rollout scoring should prioritize correctness evidence, not model confidence. A practical score can combine required-test pass rate, regression-test pass rate, static-analysis diagnostics, task-specific invariant satisfaction, visual similarity for UI work, patch size/risk, and evaluation cost. Treat hard gates (e.g., required tests) separately from soft ranking metrics; a high visual score must not compensate for a broken build. Penalize flaky or nondeterministic results and rerun selectively to estimate reliability.

When a candidate fails, convert its concrete diagnostics into the next proposer context. Avoid repeatedly exploring equivalent patches by normalizing diffs or comparing resulting tree hashes. On acceptance, produce a reviewable summary of changed files, checks run, known limitations, and artifacts. The final merge/apply action should be explicit and preserve human review controls where required.

5. HARDWARE, COMPUTE, AND RESOURCE BUDGET: TWO NVIDIA B300 GPUS

The two-GPU configuration is a planning target, not a guaranteed performance figure. Exact usable memory, bandwidth, model fit, inference rate, and concurrency depend on the specific B300 product/configuration, serving stack, quantization, sequence length, batch size, and thermal/power limits. Benchmark the chosen deployment rather than relying on theoretical peak specifications.

5.1 Memory allocation

Partition and monitor memory for three competing workloads:

Use quotas, admission control, and a scheduler. A useful initial policy is to keep one proposer model resident, limit simultaneous generations, and run cheap static checks concurrently with inference while placing browser/test jobs in bounded worker pools. Determine actual limits through load tests on representative repositories.

5.2 Throughput and search-latency economics

Measure end-to-end accepted-patch throughput, not just tokens per second. Track time for context construction, generation, patch application, LSP updates, lint/compile, targeted tests, full tests, browser rendering, and reruns. Report median and tail latency, pass rate per candidate, accepted changes per GPU-hour, and cost per accepted patch.

Use a staged evaluation funnel: cheap parsing and diff validation; incremental LSP/type/lint checks; targeted tests; broader suites; and visual comparison when needed. Stop spending on a branch as soon as a hard gate fails. Allocate a per-task wall-clock and candidate budget, then use early results to decide whether to expand search, request clarification, or return the best verified candidate.

Search economics depend on candidate yield. If each candidate is expensive and most fail late, improve context selection, static pruning, and targeted test selection before increasing branching. If candidates pass checks but vary in quality, spend additional search budget on diverse alternatives and compare them using explicit acceptance criteria. Parallelize independent candidates only while memory, CPU, and test resources remain within measured limits; contention can increase latency enough to erase the gains.

6. OPERATIONAL SAFEGUARDS AND EVALUATION

7. END-TO-END WORKFLOW

  1. 1.Parse the task into requested behavior, constraints, and acceptance checks; ask for clarification when essential requirements are ambiguous.
  2. 2.Inspect the repository, establish a clean baseline, and query the LSP for relevant definitions, references, types, and affected callers.
  3. 3.Build a bounded context package including related source, tests, configuration, and invariants.
  4. 4.Generate one or more isolated unified-diff candidates with the compact proposer.
  5. 5.Validate patch scope and apply each candidate to a disposable git state.
  6. 6.Run deterministic gates in increasing cost order; prune immediately on hard failures and feed diagnostics back into search.
  7. 7.For visual work, render desktop/mobile viewports, calculate perceptual metrics and spatial errors, and iterate targeted frontend changes.
  8. 8.Rank surviving candidates by correctness evidence, task-specific quality, risk, and compute cost; rerun checks when reliability is uncertain.
  9. 9.Present the best verified patch with exact checks and limitations. Apply or merge only through the explicitly authorized workflow.

CONCLUSION

The proposed architecture uses neural models for flexible code proposal and deterministic repository tools for grounding, constraint checking, and execution evidence. Its advantage is not that verification makes every candidate correct, nor that a compact model automatically beats a larger one. The advantage is a disciplined feedback loop: proposals are cheap enough to explore, invalid states are pruned early, and claims are tied to observable repository checks. On two B300 GPUs, resource-aware scheduling and test-time search should be optimized around accepted, verified changes per unit time - not raw model scale alone.

Download full proposal (.txt)

The proposal above matches the downloadable file verbatim.