Aug 9, 2026 Research Publication Safety
Meridian 0.1: Why Data Design Beats Method Innovation
A 135M-parameter case study in LaTeX-aware academic proofreading — curated data and task decomposition delivering a 2.2× ROUGE-L gain, scaling to Meridian-AC-Nano at 751.6M parameters.
Meridian 0.1 — Data Design × LaTeX-Awareness
135M-parameter case study: LoRA fine-tuning on a 681-pair LaTeX-aware corpus. The same recipe scales to Meridian-AC-Nano, 751.6M parameters, saturating code and LaTeX while preserving restraint and safety.
Meridian-AC-Nano
751.6M
parameters · Qwen3-0.6B base · LoRA
Reported ROUGE-L
0.641
Meridian 0.1
135M reference · measured
0.929
AC-Nano
751.6M · measured
Reported evaluation results on the held-out test sets described in the paper. Not directly comparable to unrelated models or datasets unless the paper explicitly establishes that.
Overview
Meridian 0.1: Why Data Design Beats Method Innovation: A Case Study of a LaTeX-Aware Academic Proofreading SLM
Meridian 0.1 presents a data-first approach to academic proofreading using a LaTeX-aware small language model. The work investigates how training-data design, task decomposition, and document-structure awareness can improve academic proofreading without relying exclusively on larger model architectures.
For academic proofreading to be trustworthy, a model must respect document structure — citations, equations, and LaTeX commands — not just prose. We curated 681 LaTeX-aware pairs (544 train / 66 validation / 71 test) and fine-tuned with LoRA/PEFT. The 135M-parameter reference model moved from 0.290 zero-shot to 0.641 ROUGE-L — a 2.2× gain — with 76.2% LaTeX validity on the evaluated subset.
The same recipe scales to Meridian-AC-Nano, a 751.6M-parameter multi-task academic assistant built on a Qwen3-0.6B base with 4.6M trainable parameters, reporting 0.929 ROUGE-L on its held-out sets while preserving restraint and safety behavior.
Yuvraj Pandey · Shivanshi Tripathi — Cogerphere AI Labs
August 9, 2026 · Technical Report v1.1 · Preprint / Technical Report · DOI: 10.5281/zenodo.21863555 · CC BY 4.0
Research at a glance
Research Focus
- Data-centric AI
- Small Language Models
- Academic proofreading
- LaTeX-aware processing
- Multi-task learning
- Dataset engineering
- Model evaluation
Model
Meridian-AC-Nano
751.6M parameters · 4.6M trainable (LoRA)
What Meridian proves
2.2×
ROUGE-L gain over zero-shot (0.290 → 0.641)
76.2%
LaTeX validity (subset)
681
Curated pairs (544 train / 66 val / 71 test)
Research pipeline
Data → curation → PEFT → evaluation → GGUF deployment for local-first inference. Every number in the report is labeled [M] Measured or [T] Target, and this page reports evaluation results exactly as described in the paper.
- Research Data
- Dataset Curation
- Instruction Formatting
- LoRA / PEFT Fine-Tuning
- Evaluation
- Quantization
- GGUF
- Local Inference

Meridian data design overview — see full paper for ablations and error analysis
Cite this work
Cite this work
Copy the citation or BibTeX — works on desktop and mobile.
Pandey, Y., & Tripathi, S. (2026). Meridian 0.1: Why Data Design Beats Method Innovation: A Case Study of a LaTeX-Aware Academic Proofreading SLM (Version 1.0.1). Zenodo. https://doi.org/10.5281/zenodo.21863555
BibTeX
.bib@techreport{pandey2026meridian,
author = {Pandey, Yuvraj and Tripathi, Shivanshi},
title = {Meridian 0.1: Why Data Design Beats Method Innovation: A Case Study of a LaTeX-Aware Academic Proofreading SLM},
year = {2026},
institution = {Cogerphere AI Labs},
type = {Preprint},
doi = {10.5281/zenodo.21863555},
url = {https://doi.org/10.5281/zenodo.21863555}
}Tip: On mobile, long-press the block to select manually if clipboard is blocked. The button falls back to legacy copy.
Publication type
Preprint / Technical Report
Authors
Yuvraj Pandey · Shivanshi Tripathi
Cogerphere AI Labs
Publication information
Publication information
- Title
- Meridian 0.1: Why Data Design Beats Method Innovation: A Case Study of a LaTeX-Aware Academic Proofreading SLM
- Version
- Technical Report v1.1 · Preprint / Technical Report — not peer-reviewed journal publication
- /meridian-0-1.pdf (crawlable, no auth)
Reporting policy: every number in the report is labeled [M] Measured or [T] Target. This page reports evaluation results as described in the paper and labels them as reported results.
Appendix: methods & artifacts — Publication · Aug 9, 2026 · 14 min read · Filed under Research, Publication, Safety. Every number in the report is labeled [M] Measured or [T] Target.
← All research