AI Tools

The Best AI Tools for Writing Research Papers, After Three Weeks of Actually Using Them

I did not set out to write a roundup. I had a systematic review stuck at forty percent since March, roughly 140 PDFs I had skimmed and not really read, and a supervisor asking for a full draft by month end. So I signed up for everything, put the same three questions through each tool, and kept a spreadsheet of what came back.

Three weeks later the spreadsheet said something I was not expecting. The tools that saved me real time were not the ones that write. They were the ones that read. Every hour I got back came from screening, extraction and citation checking. The drafting tools mostly produced text I spent longer fixing than I would have spent writing it badly myself and revising once.

I also learned the hard way that two of these tools will hand you references that do not exist. Not broken links. Papers never published, with plausible authors and a DOI resolving to something else entirely. That happened in week one and it changed how I used everything afterwards. So this is not a feature list copied off pricing pages. It is what held up, what quietly broke, and what I still pay for now the review is submitted.

How I judged them

ItemDetail
Test period21 days, one live manuscript
FieldHealth behaviour, plus two cross-checks in materials science
QueriesThe same 3 questions per tool, run twice
VerificationEvery returned citation checked in Crossref

Scores are mine, weighted toward one thing: how much output I could use without re-checking from scratch. A tool that finds forty papers and gets six wrong is worse than one that finds twenty and gets none wrong, because I have to verify all forty either way.

Table 1. All eight scored, pricing verified 28 July 2026

ToolBuy it forFree tier enough?AnnualMonthlyScore
ElicitScreening and pulling data out of papersCapped credits, enough to judge it$12$499.1
UndermindFinding what everything else missed3 deep searches a month$16-8.8
sciteChecking a claim survived the last decadeToo thin to be useful$12$208.6
ConsensusAnswering a question in ninety secondsResets monthly, genuinely usable$9$458.4
Gemini NotebookQuestions across your own PDF pile50 sources, most people stay hereBundled-8.3
ResearchRabbitSeeing the shape of a fieldFull product, no time limit$10-8.2
SciSpaceGetting through a paper you cannot parseTight caps, runs out fast$12$707.8
PaperpalLanguage and pre-submission checks200 edits a month, about two sessions$12$257.6

Pricing moves constantly and several vendors run region-based or student rates. Every figure was checked on 28 July 2026. Confirm on the vendor page before you buy.

If you only want to be told what to buy

The honest answer depends less on your field than on what stage of the work keeps stalling. This is the table I wish someone had handed me in week one.

Table 2. Pick by situation rather than by feature list

If this is youStart hereAdd when you canDo not bother
Undergraduate or taught mastersResearchRabbit and Gemini Notebook, both freeConsensus free tierAnything paid, for now
PhD student running a reviewElicit Plusscite for the month before submissionSciSpace, it overlaps Elicit
Postdoc or faculty, publishing oftenElicit Proscite annual, Undermind per projectDrafting tools of any kind
Clinician checking evidence fastConsensusscite for contested claimsElicit, unless you run reviews
Writing in English as a second languagePaperpal annualElicit for the reading loadGeneric grammar tools
Reading well outside your fieldSciSpaceGemini Notebook for your own setUndermind until you have scoped it

The eight tools, ranked

Elicit

BEST OVERALL   Free / ~$12 / $49 Pro   ·   ~138M papers   ·   Screening and extraction   ·   9.1 / 10
Elicit changed my workflow rather than speeding it up. You give it a question, it returns papers, then you add columns: sample size, study design, outcome measure, effect direction. It fills the grid by reading the papers. A fortnight of tabbing between PDFs and a spreadsheet became about a day and a half of checking a grid it had already populated.

Introducing Elicit Alerts - Elicit

Roughly one cell in ten needed correcting, usually where a paper buried something in a supplementary file. But correcting a wrong cell takes thirty seconds. Building the row from nothing takes twenty minutes.

HELD UP

• Custom extraction columns, best in category

• Finds papers using different vocabulary for the same idea

• Clean CSV and BibTeX export for PRISMA workflows

GAVE ME TROUBLE

• Quality drops when only the abstract is accessible

• Weak on theoretical work with no measurable outcomes

• Big jump from the $12 tier to the $49 tier

Bottom line. If you are doing a literature review with any structure to it, this is the one to buy first. Everything else on this list is optional in a way Elicit is not.

Undermind

BEST FOR HARD QUESTIONS   Free / ~$16 a month annual   ·   Agentic citation-trail search   ·   Niche and cross-field topics   ·   8.8 / 10
Undermind is slow on purpose. You describe a topic, it disappears for several minutes, reads full texts where it can, follows citation trails, and returns a report plus a coverage estimate of how thoroughly it thinks it searched. That number is why I kept it. Nothing else here tells you when you have probably run out of relevant literature.

Undermind.ai platform has been revamped with the Undermind projects feature  | Singapore Management University (SMU)

It found four papers I had missed after two months of searching, all in adjacent fields using different terminology. That is the job it does. Ask it something broad and well covered and it will underwhelm you.

HELD UP

• Surfaces genuinely obscure work in neighbouring fields

• Coverage estimate gives a defensible stopping point

• Explains why each paper was retrieved

GAVE ME TROUBLE

• Each search takes minutes, so no casual browsing

• Free tier covers evaluation and not much more

• Poor first stop for broad, unscoped topics

Bottom line. Buy it for the last twenty percent of a search, once the obvious papers are already in your library and you suspect something is missing.

scite

BEST FOR CITATION CHECKING   $20 a month, ~$12 annual   ·   1.2B+ citation statements   ·   Testing whether a claim holds   ·   8.6 / 10
scite does not count citations, it classifies them. For any paper you get how many later papers supported the finding, contrasted it, or merely mentioned it, with the sentence from each citing paper. Citation counts tell you a paper was noticed. This tells you whether it survived.

SciTE Lua Scripting Extension

Reference Check earned its keep. Upload your manuscript and it flags references carrying editorial notices, retractions, or a real body of contrasting citations. It caught a paper in my list that had picked up a correction I had not seen.

HELD UP

•  Supporting versus contrasting counts change how you read

•  Catches retractions before a reviewer does

•  Extension surfaces the data on journal sites

GAVE ME TROUBLE

•  Most citations classify as mentioning, accurate but unhelpful

•  Contrasting citations are rare outside contested areas

•  Humanities coverage is thinner than the sciences

Bottom line. Not a discovery tool. Buy it for the fortnight before submission and run your whole reference list through it.

Consensus

BEST FREE TIER   Free / ~$9 to $12   ·   200M+ papers   ·   Quick evidence checks   ·   8.4 / 10
Consensus answers yes-or-no research questions and shows a meter of how the retrieved studies split. It is the tool I opened most often and thought about least, which sounds like an insult and is not. When you need to know in ninety seconds whether a claim is broadly supported before building a paragraph on it, nothing else is close.

Introducing: Consensus 2.0 - Consensus: AI Search Engine for Research

The meter summarises the top retrieved studies. It is not a meta-analysis, and that distinction matters. It also refreshes its free allowance monthly instead of handing you a one-time credit pool, which makes it the most usable free tier here.

HELD UP

• Every claim links straight back to its paper

• Free allowance resets monthly, so it stays usable

• Exports to Zotero, Mendeley and EndNote on all tiers

GAVE ME TROUBLE

• The meter can imply more agreement than the studies support

• Struggles with anything not phrased as a testable claim

• Not exhaustive, so never use it alone for a review

Bottom line. Start free, stay free for a month, and only upgrade if you find yourself hitting the analysis cap every week.

Gemini Notebook (formerly NotebookLM)

BEST FOR YOUR OWN PDFS   Free / Google AI plans   ·   50 sources per notebook free   ·   Synthesis you can trace   ·   8.3 / 10
Google renamed NotebookLM to Gemini Notebook on 16 July 2026. Same product, same site, new logo, so ignore anything saying it was discontinued. What matters is that it answers only from documents you upload, with an inline citation on every sentence that jumps to the exact passage. When sources are fixed, that constraint removes the fabrication problem almost entirely.

How to Use NotebookLM with Gemini in 2026 (with 3 Practical Use Cases) | by  Uzman Ali | Write A Catalyst | Medium

I used it as a reading assistant, not a search tool. Twenty-five papers into one notebook, then questions like which of these used an intention-to-treat analysis. The free plan allows fifty sources per notebook, which is the wall most people hit first.

HELD UP

• Answers grounded in your uploads, with traceable citations

• Free tier is unusually generous

• Audio overviews are good for reviewing on a commute

GAVE ME TROUBLE

• No cross-notebook search, so large libraries fragment

• No formatted citation export

• Check policy before uploading unpublished data

Bottom line. The best free tool in this article, provided your project fits inside one notebook. It stops being useful the moment it does not.

ResearchRabbit

BEST FOR MAPPING A FIELD   Free / RR+ ~$10 a month   ·   310M+ articles   ·   Snowballing and lineage   ·   8.2 / 10
Feed it a few papers you already trust and it draws the citation network around them, time on one axis and influence on the other. Within ten minutes you can see what everything descends from, where the clusters sit, and which corner of the field you have been ignoring. It is the fastest way I know to orient yourself somewhere unfamiliar.

ResearchRabbit: AI Tool for Smarter, Faster Literature Reviews

The Zotero sync is the quiet reason it stays. Your library appears as collections, new finds save back in one click, and the export and import shuffle disappears. It was free for years and has since added a paid tier, but the free version still covers what most individual researchers need.

HELD UP

• Graphs make gaps visible in a way lists never do

• Two-way Zotero sync, smoothest integration here

• Suggestions improve as your collection grows

GAVE ME TROUBLE

• No AI summarisation, so it pairs with Elicit

• Networks get crowded past a few hundred nodes

• Output depends heavily on your seed papers

Bottom line. Free, fast, and the first thing I open when starting in a field I do not know. There is no reason not to have an account.

SciSpace

BEST FOR DENSE PAPERS   Free / ~$12 to $20   ·   280M+ papers   ·   Reading outside your field   ·   7.8 / 10
SciSpace tries to be everything: search, PDF chat, writing, journal formatting, patent search. Everything-tools usually disappoint and this one mostly does not, though the parts are uneven. Chat with PDF is the strong piece. Highlight a paragraph of statistics you cannot follow, ask what it means, and it explains with the passage alongside so you can check it.

Introducing SciSpace's AI-powered literature review

Deep Review sits on the top tier, priced well above the standard plan and a real commitment. For most people the mid tier is the sensible stopping point. Treat the writing tools as a bonus, not a reason to subscribe.

HELD UP

• Handles methods, tables and equations well

• Useful when reading outside your specialism

• Extension works on Google Scholar and PubMed

GAVE ME TROUBLE

• Deep Review sits on a tier most will not justify

• Abstract-only access quietly weakens some answers

• Many jobs, none of them best in class

Bottom line. Worth it if reading is your bottleneck rather than finding. If you already have Elicit, the overlap is real.

Paperpal

BEST BEFORE SUBMISSION   Free / $25 a month, $139 a year   ·   200 free edits a month   ·   Language and journal checks   ·   7.6 / 10
Paperpal comes out of Cactus Communications, which has been doing academic editing for two decades, and it shows. It corrects toward scientific register rather than general readability. Grammarly will happily flatten a hedged claim into a confident one. Paperpal mostly leaves hedges alone, because in academic writing the hedge is the point.

AI Academic Writing Tool - Online English Language Check | Paperpal by  Editage

Its submission checks cover structure, reference formatting and technical compliance. They do not tell you whether the science is ready, whether the journal is a sensible target, or whether a reviewer will find a hole. Do not mistake one for the other.

HELD UP

• Best academic tone correction I tested

• Runs in Word, Google Docs and Overleaf at no extra cost

• Annual pricing is under half the monthly rate

GAVE ME TROUBLE

• Heavy rewrites soften novelty claims into safe phrasing

• Free tier lasts about two serious editing sessions

• Submission checks are technical, not scientific

Bottom line. Accept sentence-level suggestions, reject paragraph-level rewrites, and you will get most of the value.

Two I stopped paying for

Jenni AI does what it says: inline autocomplete, citations inserted as you write, thousands of styles. It broke my writing rather than helped it. Accepting a suggested sentence is easy, and after twenty of them the paragraph argued something slightly different from what I meant. The auto-suggested citations needed checking every time, which cancelled out the speed. If autocomplete suits how you think, the annual plan is cheap. It did not suit mine.

General chatbots for drafting. I still use them daily for reasoning out loud, restructuring an argument, and explaining a method three different ways. Not for text that goes into the manuscript, and never for finding citations. That is where the fabricated references came from.

Why you verify every reference, every time

This is not caution for its own sake. It is the best documented failure mode in AI-assisted research writing, and the numbers are worse than most people assume.

FigureWhat it measured
55%of GPT-3.5 references in a 2023 Scientific Reports study were fabricated outright, against 18% for GPT-4, across 636 citations
43%of the GPT-3.5 references that were real still contained substantive errors in volume, issue or page details
28.6%hallucination rate for GPT-4 on systematic review prompts in a 2024 JMIR analysis of 471 references
38%DOI validity in humanities prompts in one cross-model audit, against 71% for the natural sciences

The pattern underneath is the useful part. Fabrication tracks how often a topic appears in training data. One GPT-4o study found roughly 6% fabrication on a heavily studied condition like major depressive disorder, rising to around 28% for less studied ones. The model invents sources precisely when you are working on something original, which is exactly when you are least equipped to notice.

Tools that retrieve from a real index, then cite what they retrieved, do not have this failure mode in the same way. Tools that generate a citation from memory do.

Elicit, Consensus, scite and Undermind search first and answer second. A general chatbot answers first. It is the difference between a librarian and a very confident person who has read a lot.

What journals actually require you to disclose

Most tool roundups skip this, and it is the part that can cost you a publication. The position across standards bodies settled in 2023 and has tightened since. Every body below prohibits listing AI as an author, so the table covers what actually varies.

Table 3. Where the disclosure goes, and what each one catches you on

Publisher or bodyWhere the statement goesThe part people miss
ICMJECover letter and manuscript, bothWriting help goes in acknowledgements. Data collection or analysis goes in methods. Two different places.
COPEPer the journal's own formatPosition covers reviewers and editors too, not only authors.
ElsevierIts own section, before the referencesPolicy updated September 2025. AI-generated images are not permitted at all.
WileyMethods, or a disclosure or acknowledgements sectionWhich section depends on what the tool actually touched.
SAGEOnly when the AI generated somethingRefining your own wording counts as assistive and needs no disclosure. Generating new text does.
Science (AAAS)Detailed reporting in the manuscriptStrictest of the set. Expects tool names and versions, sometimes the prompts.

Three rules run through all of it. Software cannot be an author, because it cannot be held accountable. You remain responsible for every word, including anything a tool produced. And AI output is never a citable source, so a reference must point to a real, verifiable publication.

Check your specific target journal, not just the publisher. Requirements vary between titles in the same house, and the ICMJE recommendations were revised in January 2026 with an expanded section covering AI across the whole publishing workflow. If you review for a journal, most policies now prohibit putting someone else’s manuscript into a general AI tool at all, on confidentiality grounds.

The order I would use these in

Table 4. Eight stages, and what each one costs you

StageWhat I openCostWhat you get out of it
1. Orient yourselfResearchRabbitFreeThree seed papers, read the map, find the clusters you did not know about.
2. Search properlyElicit, UndermindPaidElicit for structured retrieval. Undermind when you suspect something is still missing.
3. Screen and extractElicitPaidExtraction tables. This is the step where the days come back.
4. Read the hard onesGemini Notebook, SciSpaceMixedNotebook for your own PDF set. SciSpace when a single paper is defeating you.
5. Sanity check claimsConsensus, sciteMixedConsensus for speed. scite for whether a finding actually held up.
6. Write itNothingFreeNo tool. This is the part where the thinking happens.
7. Audit and polishscite, PaperpalPaidReference Check across the whole bibliography, then Paperpal for language and compliance.
8. DiscloseNothingFreeWrite the statement your target journal requires before you submit, not after a reviewer asks.

What I still pay for

The review went in on the nineteenth. Here is what survived the trial: Elicit annually, scite for a month either side of submission, and free accounts on ResearchRabbit, Consensus and Gemini Notebook that I use constantly and have never paid for. I cancelled the rest. Undermind I will resubscribe to when the next project starts, because it earns its money in bursts rather than continuously.

If you want one recommendation and nothing else: buy Elicit, keep ResearchRabbit and Gemini Notebook open in other tabs, and verify every reference before it enters your bibliography. That costs about $12 a month and did roughly eighty percent of what all eight tools did together.

What I keep coming back to is how little of this was about writing. I went looking for tools to help me write a paper and came away with tools that help me read faster, search wider and check harder. The writing stayed where it was, which is me, a document, and an uncomfortable number of afternoons. I think that is right. The argument is the paper. Everything else is logistics, and logistics is what these things are good at.

The real change is that I no longer trust my own reading of a field after a week in it. Undermind found four papers I had missed. scite showed me a correction notice on a source I had cited without checking. Both were my failures, not the literature’s, and I would not have caught either one alone.

Related Posts