Skip to content

Latest commit

 

History

History
545 lines (351 loc) · 20.3 KB

File metadata and controls

545 lines (351 loc) · 20.3 KB

The Research Operating System

The Ultimate Academic Research Application — Complete Vision


Part I — The Concept

The ultimate academic research tool is not a better reference manager, PDF reader, AI chatbot, or note-taking app.

It is a Research Operating System.

Its purpose is to take a world-class researcher from:

curiosity → question → evidence → argument → method → analysis → writing → review → publication → impact

without losing provenance, rigour, intellectual control, or the researcher's own thinking process.

The Two Core Principles

1. The core object is not the document. The core object is the research question — and beneath it, the claim.

Everything flows from the question. Everything is built from claims.

2. Research is not a waterfall. It is a loop.

Analysis changes the question. Reading kills the hypothesis. Writing exposes gaps in evidence. The system must treat backward movement as normal, sanctioned, and recorded — not as failure.

The corrected flow:

Question ⇄ Evidence ⇄ Claims ⇄ Argument — with writing happening continuously, and every loop versioned in a Decision Log.

Founding Principles (Non-Negotiable)

  1. No fake citations. AI may never invent a reference.
  2. No unsupported claims. Every claim knows its evidence.
  3. No broken provenance. Every note knows where it came from.
  4. No black-box outputs. Every result links to data, method, and decision.
  5. No lock-in. Open formats everywhere. Local-first storage. Your five years of PhD work is yours.
  6. The researcher remains the author. AI is the assistant, never the scholar.
  7. Ethics is a thread, not a stage. Sensitivity and consent travel with every data item.

Part II — The Flow

Stage 0. Quick Capture Inbox (new)

Ideas arrive in the shower, in the car, in the archive basement — not in the Reading Room.

The system provides frictionless capture:

  • mobile app with one-tap note
  • voice memo with automatic transcription
  • email-in address per project
  • browser clipper
  • photo of a physical document or whiteboard
  • forward a WhatsApp/Signal message

Everything lands in an Inbox with timestamp and origin. Nothing is lost. Triage into the knowledge graph happens later, on the researcher's terms.

The inbox is the front door of the research mind.


Stage 1. Research Intent

The researcher creates a Research Mission, capturing:

  • research field
  • broad topic
  • problem area
  • initial hypothesis
  • key concepts
  • preferred theories
  • possible methods
  • intended outputs
  • target journals or conferences
  • ethical sensitivity
  • expected datasets or archives
  • collaborators
  • funding or deadline context

The system creates a Research Command Centre around this mission.

The first question is always:

What are you trying to prove, explain, question, challenge, or discover?


Stage 2. Question Builder

The system refines the research question before deep source collection:

  • broad topic → problem statement → research gap
  • primary research question
  • secondary questions
  • hypothesis or proposition
  • scope boundaries and exclusions
  • key definitions
  • assumptions
  • risks of bias

It diagnoses whether the question is:

  • too broad / too narrow
  • already answered
  • methodologically weak
  • ethically sensitive
  • misaligned with available evidence
  • likely publishable

Output: a structured Research Design Brief — versioned, so when the question evolves (it will), every version is retained with the reason for change.


Stage 3. Field Mapping — Living, Not Static (upgraded)

Before reading deeply, the tool maps the field as a knowledge landscape, not a search results page:

  • key authors and schools of thought
  • landmark publications
  • recent debates and contradictions
  • dominant and neglected theories
  • methodological traditions
  • journals, conferences, datasets, archives
  • citation networks: who cites whom, which ideas cluster

The upgrade: the map never sleeps.

Connected continuously to open scholarly indexes (OpenAlex, Semantic Scholar, Crossref, CORE), the system runs Field Alerts:

  • new papers touching your claims
  • retraction warnings on any source you cite
  • new contradictions entering the debate
  • a rising author or dataset in your niche
  • a competing paper answering your question

Research questions rot. The map keeps them fresh — for the entire multi-year life of a PhD.


Stage 4. Evidence Intake

Ingestion from every evidence type:

  • academic databases, journal articles, books, theses, conference papers
  • datasets, archival records, field notes
  • interviews, images, audio, video
  • webpages, government reports, legal and policy documents
  • lab results, survey outputs, code repositories

Each item enters with full metadata: author, title, date, source, DOI/URL, citation data, document type, reliability level, access rights, version, notes, and relationship to the research question.

Every source is traceable from the moment it enters.

There are no orphan PDFs.


Stage 5. Source Triage

The researcher is never forced to read everything immediately.

Triage categories: essential · useful · background · contested · weak · duplicate · excluded · read later · method source · theory source · evidence source.

For each source, an AI-generated structured preview:

  • main argument, methodology, evidence, key findings
  • limitations and possible bias
  • relevance to the research question
  • whether it supports or challenges the hypothesis
  • cited authors, useful quotations, extracted concepts

Every preview is labelled:

AI preview — not human verified.

The system never pretends the researcher has read something. Read-status is tracked honestly: unread / previewed / skimmed / read / deeply read.


Stage 6. Deep Reading Room

Serious scholarly reading:

  • PDF reading, highlighting, margin notes
  • concept / claim / method / theory / contradiction / uncertainty tagging
  • quote extraction with page-level references
  • side-by-side comparison of sources
  • AI explanation of difficult passages
  • translation
  • summaries at article, section, and paragraph level

Every highlight asks one question:

What role does this play in your research?

Roles: supports my argument · challenges my argument · defines a concept · provides evidence · provides method · provides theory · historical context · shows a gap · shows a contradiction · useful quote · weak claim · needs verification.

This turns reading into structured research, not passive highlighting.


Stage 7. Research Knowledge Graph

The heart of the system. A graph — built automatically and manually — of:

concepts · authors · theories · methods · datasets · sources · claims · evidence · counterarguments · places · institutions · time periods · citations · quotations · notes · chapters · article sections.

The graph answers:

  • Which sources support this claim? Which contradict it?
  • Which theory is linked to this method?
  • Which authors dominate this debate?
  • Which evidence is weak? Which claims have no citation?
  • Where is my argument overdependent on one source?
  • What have I not read yet? What am I assuming?

The application does not only store research.

It reveals the structure of the researcher's thinking.


Stage 8. Claim Ledger — The Spine

Every academic output is built from claims. This is the killer feature and the first thing to build.

Each claim carries:

  • claim text and status
  • supporting sources / opposing sources
  • confidence level
  • evidence type
  • relevant quotations with page references
  • methodological basis and theory connection
  • researcher notes
  • unresolved weaknesses
  • ethical concerns
  • whether it is original, derived, or speculative

Claim statuses: idea → working claim → supported / contested / weak / rejected → needs more evidence → publishable.

Before the researcher writes a paragraph, the system already knows what evidence stands behind it.

A normal tool manages documents. This tool manages claims and evidence.


Stage 9. Decision Log — The Memory of Every Loop (new)

Examiners and reviewers always ask: "Why did you exclude X? Why did the scope change? Why was that case dropped?"

The Decision Log records, automatically and manually:

  • every scope change, with date and reason
  • every exclusion of a source, case, or dataset
  • every rejected or revised hypothesis
  • every methodological pivot
  • every question reformulation
  • every supervisor instruction acted upon

This is the audit trail of thinking. It:

  • answers examiner questions with receipts
  • feeds the limitations section for free
  • protects against accusations of cherry-picking
  • preserves intellectual honesty across years
  • makes the iteration loop safe — going backwards is recorded, not erased

Waterfall tools punish changing your mind. This system versions it.


Stage 10. Method Design Studio (trimmed and sharpened)

Method support ships as discipline templates, not as an attempt to natively encode every methodology:

  • design science · archival method · ethnography · case study
  • qualitative / quantitative / mixed methods
  • discourse analysis · historical method · legal-policy analysis
  • computational and digital humanities methods

Each template guides: research design, sampling, data sources, instruments, coding framework, variables, validity, reliability, ethics, consent, bias control, reproducibility, and data management plan.

Output: a Method Protocol that flows directly into the thesis methodology chapter, the journal article, the grant proposal, and the ethics application — written once, reused everywhere.


Stage 11. Analysis Bridge (redesigned — integrate, don't rebuild)

Rebuilding NVivo, SPSS, and Jupyter inside one application is scope death. The value is provenance, not the analysis engine.

The system therefore acts as a bridge:

  • launch external notebooks (Jupyter, R, Python), QDA tools, statistics packages
  • light built-in support for the basics: thematic coding, memo writing, simple visualisation, transcript annotation
  • ingest every result back with full provenance metadata:
    • source data and version
    • method and code used
    • date generated
    • researcher decision attached
    • link to the claim(s) it supports or weakens

Every chart, table, theme, and statistic in the final thesis traces back to its origin in one click.

No black-box outputs. Ever.


Stage 12. Argument Builder

Where the intellectual contribution takes shape:

  • central thesis
  • chapter and article structure
  • argument chain and evidence map
  • counterargument map
  • theory–method–evidence alignment
  • novelty and contribution statements

The researcher drags claims into an argument sequence:

  1. Problem → 2. Gap → 3. Theoretical frame → 4. Method → 5. Evidence → 6. Analysis → 7. Counterargument → 8. Contribution → 9. Implication

The system warns:

  • this claim lacks evidence
  • this section over-cites one author
  • this argument skips a step
  • this source is outdated or retracted
  • this contradicts another claim
  • this paragraph needs a primary source
  • this conclusion is stronger than the evidence allows

Stage 13. Writing Studio — Write As You Go (upgraded)

Writing is not Stage 13 in time — it is continuous. Memos become paragraphs; paragraphs become chapters. The Writing Studio is open from day one.

The writing environment is connected to the graph, never a blank page:

  • claim ledger, sources, quotes, notes, citations at hand
  • outline and argument map alongside
  • journal requirements and supervisor comments in context
  • full version history

Supported outputs: thesis · journal articles · grant proposals · conference papers · literature reviews · systematic reviews · methodology sections · policy briefs · book chapters · public summaries.

AI drafting under strict control:

  • no uncited claims
  • no invented references
  • all AI-generated text labelled
  • citation suggestions link only to real, ingested sources
  • paraphrases preserve meaning; quotes remain exact
  • researcher approval required for everything

Format freedom: native Markdown and LaTeX, export to Word/ODT, citation style switching at export, Zotero/BibTeX/RIS round-tripping.


Stage 14. Review Studio (merged: supervisor review + peer review simulation)

One studio for all review, internal and simulated.

Supervisor and co-author workflow:

  • comment threads anchored to claims, not just text
  • revision tracking and response history
  • task assignment per chapter or section

Adversarial review simulation — a hostile but fair reviewer testing:

  • originality, clarity, evidence strength
  • methodology and theory alignment
  • structure, contribution, literature coverage
  • citation quality, ethical issues, overclaiming
  • weak transitions, unsupported conclusions, journal fit

Generating reviewer-style feedback: major concerns · minor concerns · likely objections · required revisions · rejection risks · strongest contribution · weakest section · missing literature.

The simulator can adopt personas: the methodologist, the theory purist, the statistician, Reviewer 2.


Stage 15. Publication Studio (focused)

The essentials, done well:

  • journal and conference matching (scope, impact, open-access options, turnaround time)
  • submission requirement compliance: word count, formatting, reference style, data availability statement, ethics statement
  • response-to-reviewers workflow with full revision history
  • version tracking: submitted → reviewed → revised → accepted → published
  • DOI, repository deposit, open data and code deposit, public summary

Research does not end at writing. It ends when the work enters the scholarly record properly.


Stage 16. Research Memory and Future Work

After each project, the system retains the researcher's intellectual memory:

  • unresolved questions and rejected ideas (from the Decision Log — free of charge)
  • future article opportunities
  • unused sources and abandoned hypotheses
  • datasets for reuse
  • possible collaborations, conference ideas, grant opportunities
  • teaching and public communication material

The next project starts smarter because the previous one is not forgotten.

A career becomes a connected graph, not a folder of dead projects.


Part III — Dream Big: The Moonshot Features

17. The Contradiction Engine

The system continuously scans the entire claim ledger and knowledge graph for:

  • claims that contradict each other across chapters
  • a source cited as support in Chapter 2 and dismissed in Chapter 5
  • evidence that weakened after a new paper arrived
  • definitions that drifted between the proposal and the thesis

No human can hold a 90,000-word thesis in working memory. The machine can.

18. The Provenance Chain — Cryptographic Research Integrity

Every claim, quote, dataset, and analysis result carries a tamper-evident provenance chain: who added it, when, from where, and every transformation since.

When research integrity is questioned — and in the AI era it will be — the researcher can prove, cryptographically, that the work is theirs and the evidence trail is intact.

This becomes the defence dossier: a one-click export showing every claim, its evidence, its decisions, and its history. Walk into the viva untouchable.

19. The Time Machine

Scrub backwards through the project. See the knowledge graph as it stood in month 3 versus month 30. Watch the question evolve. Show an examiner exactly how the thinking developed.

This is also the honesty engine: it makes retrofitting a hypothesis to the results visibly impossible.

20. The Examiner / Reviewer Twin

Feed in the actual examination criteria, the target journal's review guidelines, or a known examiner's published work. The system builds a review twin that critiques the manuscript from that specific perspective.

Not generic feedback — targeted pre-emption.

21. Cross-Researcher Federation (Opt-In)

Researchers can selectively federate parts of their graphs:

  • a supervisor sees the claim ledger, not the messy notes
  • co-authors share an evidence pool with merged provenance
  • a research group builds a shared field map that compounds across cohorts
  • a department's collective memory survives staff turnover

Privacy-first: nothing leaves the local store without explicit, granular consent.

22. The Replication Pack

One click produces everything needed for replication: method protocol, data (where ethics permit), code, decision log, analysis provenance. The reproducibility crisis is solved at the level of tooling, not exhortation.

23. Voice-First Field Mode

In the archive, the lab, or the field: speak observations, photograph documents, tag a location, and have everything land in the inbox already structured — source, place, time, project, preliminary tags. Offline-capable, syncing later.

24. The Grant Engine

The Research Mission, Method Protocol, field map, and claim ledger already contain 80% of a grant proposal. The system assembles funder-specific drafts (NRF, ERC, NIH, Wellcome formats), tracks calls matching the research profile, and recycles unfunded proposals intelligently.

25. Impact Tracking

After publication: citation alerts, policy document mentions, media coverage, syllabus appearances, dataset reuse. The system closes the loop from curiosity to demonstrable impact — the evidence base for the next promotion, rating application, or grant.


Part IV — The Architecture of Trust

Open Formats, Local First

  • storage: plain files + SQLite, human-readable, on the researcher's machine
  • sync: optional, end-to-end encrypted
  • import/export: BibTeX, RIS, CSL-JSON, Markdown, LaTeX, DOCX, PDF/A for archival copies
  • integrations: Zotero, OpenAlex, Crossref, ORCID, institutional repositories, Git
  • the exit door is always open — full export at any moment, no degradation

A research OS that holds five years of PhD work hostage is a liability, not a tool. This is a founding principle, not a feature.

AI Containment Rules

  • every AI output labelled and attributable
  • AI may only cite sources that exist in the project
  • AI previews are never marked as "read"
  • AI suggestions require explicit acceptance
  • a full AI-interaction log exists for disclosure statements (journals increasingly require this)
  • one-click AI Disclosure Statement generation for any output

Ethics as a Thread

  • consent status and sensitivity flags attached to every data item at intake
  • flags propagate: a quote from a sensitive interview carries its restrictions into every chapter that uses it
  • publication-time check: nothing under embargo or restricted consent can leave without a warning
  • ethics application generated from the Method Protocol; clearance number and date attach to the project and flow into every output automatically

Part V — The Flow in One Line

Capture → Question → living field map → intake → triage → deep reading → knowledge graph → Claim Ledger → Decision Log → method template → analysis bridge → loop back as needed → argument → write-as-you-go → review → publish → impact → memory.

With the Claim Ledger as the spine and the Decision Log as the memory of every loop.


Part VI — Build Order

Priority Feature Why
1 Claim Ledger Forces the whole system to become serious. Everything improves once claims link to evidence.
2 Decision Log Makes iteration safe and auditable. Cheap to build, enormous examiner value.
3 Open-format storage & export Trust foundation. Must exist before anyone commits years of work.
4 Quick Capture Inbox Lowest friction, highest daily-use habit former.
5 Reading Room + role-tagged highlights Turns reading into structured input for the ledger.
6 Knowledge Graph Emerges naturally once claims, sources, and tags exist.
7 Writing Studio Connected writing on top of the ledger.
8 Living Field Alerts Retraction warnings alone justify it.
9 Review Studio Simulated and supervisor review.
10 Everything else In order of the researcher's pain.

The Difference

A normal tool manages documents.

A world-class tool manages claims and evidence.

The ultimate tool manages the evolution of a researcher's thinking — with proof.

That is the difference.