AI for Researchers: Building a Workflow for Working with Sources
Five open PDF tabs, three articles bookmarked for later, and not one coherent note about what's actually in them. Here's the workflow: from organizing sources to synthesizing information into one accurate piece of writing.
Why reading isn't the same as working with sources
Five open PDF tabs, three articles bookmarked "read later," and not one coherent note about what's actually in them. If that's you, it's not laziness. Reading and working with sources are two different processes that quietly substitute for each other without you noticing.
Reading is consumption. Working with sources runs differently: extraction, comparison, and verification. Out of ten articles on a topic, you genuinely need three key ideas, and the rest is just context that either confirms or refutes those three ideas. Without that distinction, a person drowns in volume. The more sources you have open, the less you actually absorb, because all your energy goes into just wading through them rather than figuring out what actually matters.
In How to Learn with AI, I already covered the technique for working through a single source and learning a topic sequentially from a set curriculum. This is a different task: a pile of scattered material you have to pull into a coherent picture yourself, not a course with a structure already built for you. This is the work of a researcher, an analyst, a student writing a term paper. AI changes this more than it changes sequential learning, because the bottleneck here is synthesis across sources, not understanding any single source on its own. Synthesis is exactly what human memory is poorly suited for, and exactly what a language model, which can hold dozens of documents in context at once, is good at.
A practical consequence of this difference: the techniques from the piece on learning with language models, tutor mode, the Feynman technique, spaced repetition, don't carry over here one for one. They're built for a situation where the structure of the material is already set by someone else: a textbook, a course, an instructor. Working with sources starts from a situation where no structure exists yet, and the researcher's first job is discovering that structure themselves inside a pile of scattered material, before it even becomes clear what to study in what order.
What follows is the concrete process, step by step: how to organize sources, speed up reading without losing quality, synthesize conclusions from material that contradicts itself, and not lose accuracy where the cost of a mistake is high, in a thesis, a client report, or a piece of writing headed for publication.
Before getting into the concrete steps, it's worth naming the symptom that makes it easiest to spot that the process is broken right now: if after a week of working with sources you can't state in a couple of sentences what you actually learned that's new, reading happened, but working with sources didn't. The material passed in front of your eyes without turning into a single usable claim. That feeling is familiar to almost anyone who reads a lot for work or school, and it's usually the clearest signal that somewhere in the chain, the substitution described above has taken place.
What the workflow is made of
Before getting into tools, it's worth explicitly breaking the process down into steps. Most people do this unconsciously and lose information somewhere in the middle without even noticing which step it happened at.
Capture the source → Extract key claims → Verify →
Connect to other sources → Synthesize into your own writing
Every break in this chain costs time. Capture a source but don't extract its claims: a week later you'll have to reread it from scratch, because at best you'll recall a general impression, not the specific arguments. Extract the claims but don't verify them: you risk building a conclusion on someone else's factual error, or your own AI hallucination, which looks convincing right up until someone asks for the source. Verify but don't connect to the rest of your sources: you end up with a pile of disconnected facts instead of an argument, and disconnected facts read on the page exactly like a list, not like a thought.
AI is useful at every one of these steps, but differently at each one, and with a different degree of autonomy. At capture and extraction, you can trust the model with a lot of routine work. At verification, AI autonomy should be at its lowest and your own oversight at its highest. At synthesis, the model should propose structure, not write the final text in your place. Each step, covered separately below.
How AI speeds up a literature review
The literature review is a classic pain point: you need to understand what's already known about a topic before you can state your own question or hypothesis, and the volume of material is often too much to read every article in full by hand.
A working sequence that actually saves time, not just one that looks faster:
- Gather 15-30 relevant sources: search, references from articles you've already found, recommendations from your advisor or colleagues.
- Run each one through AI with a prompt like "what question does the author ask, what answer do they give, what limitations do they acknowledge themselves," not just "summarize this."
- Group the results by position: who agrees with whom, who contradicts whom, where's a genuine gap nobody has addressed.
- Write the review following this map of positions, not the chronology of publication. That immediately shows where your own contribution can fit.
The third step is the most important and the most commonly skipped. Without it, a literature review turns into a recap of ten articles in the order you happened to read them, rather than an analysis of the field as a whole. The difference shows immediately: a recap-review reads like "Smith writes X, Patel writes Y, Garcia writes Z." An analysis-review reads like "on question X there are two camps, and the disagreement comes down to differing methodology, while question Y hasn't been studied at all."
Worth a note on scale. 15-30 sources is a reasonable target for a term paper or a work report, not a universal number that fits any scope of work. For a dissertation or a serious analytical review, the count can run into the hundreds, and at that scale, manually grouping by position becomes physically unmanageable without AI help at the initial sorting stage.
NotebookLM for researchers: how it actually works in practice
NotebookLM got a passing mention in the earlier overview of learning tools. It deserves a closer look here, because for working with sources, it's arguably the most purpose-built tool on the market for exactly this job.
The key difference from an ordinary chat with ChatGPT or Claude: NotebookLM answers strictly from the sources you upload, not from the model's general knowledge. Upload 20 PDFs, and every answer comes with a precise citation pointing to a specific document and location in it, instead of a vague "research shows." This sharply cuts the risk of the model blending in something from its general training corpus and passing it off as content from your sources.
Practically useful features for working with sources:
- Grounded answers with citations. A clickable link to a specific paragraph in a specific document for every claim. You can verify it in seconds without rereading the entire file.
- Audio overviews of your sources. Generates a podcast-style discussion of your material. Surprisingly good for refreshing your memory of a large body of sources on the go (commuting, running), rather than sitting down to reread.
- Cross-document search. "Where across all these articles is X mentioned" runs against your whole source set at once, rather than one file at a time, which is what you'd have to do manually with Ctrl+F through each PDF in turn.
- Notes inside the notebook. You can jot down your own thoughts right next to the sources without switching to a separate app. This removes exactly the extra friction points in capture I wrote about in the piece on second brains.
The main limitation: NotebookLM answers poorly to questions outside the uploaded sources, and it's not built for that. Ask it something that isn't in the documents, and you get an honest "this isn't mentioned in the sources" instead of a made-up answer, which is actually a plus, not a minus, you just need to keep in mind that this is a tool for a specific body of documents, not for general knowledge. If you need open-ended conversation drawing on a model's general knowledge, that's what ChatGPT and Claude are for, and keeping both types of tool around in parallel is normal practice, not redundancy.
A practical scenario where this boundary shows up clearly: trying to ask NotebookLM about an adjacent topic that's logically connected to the uploaded sources but not directly mentioned in them. For example, if all the uploaded articles are about spaced repetition and the question is about working memory in general, the tool will honestly decline to build that bridge on its own, even if the bridge seems obvious to a person with general knowledge of the topic. This is a deliberate limitation, not a shortcoming: the developers deliberately chose strictness over a willingness to guess, because for research work, a false but plausible bridge is more dangerous than an honest refusal to answer.
How to organize sources: from PDFs to citations
Before AI can analyze anything, sources need to be organized. Otherwise even the smartest tool will dig through the same chaos you would, just with a more confident tone.
A minimal working structure that needs no special software:
- One shared folder for all PDFs and material for the project, with clear file names: author, year, a short topic label. Not "document (34).pdf," which won't mean anything even to you a month from now.
- One registry file, where AI (or you by hand) logs, for every source: the full citation, three key claims, and verification status.
- Explicit flagging of unverified claims. If AI extracted a claim but you haven't checked it against the original, that needs to be visible at a glance in the registry, not lost among verified data.
An example registry entry:
Chen, 2024, "The Effect of Spaced Repetition in Distance Learning"
Claims: (1) a 3-7 day interval gives maximum retention benefit,
(2) the effect is weaker for procedural skills than for facts,
(3) not tested on groups over 60.
Status: claim (1) verified against the original; (2) and (3) from AI, not checked.
Notice the last line: without it, a month from now there's no way to remember which of the three claims you personally verified and which you took on faith from the model. That's the difference between a working registry and just a nicely formatted list.
AI speeds up filling in the registry, but it doesn't replace it. Without one single place where information about all your sources accumulates, you're just moving the chaos from a folder of PDFs into a chat history with a bot, and finding a specific claim a month later will be just as hard as before.
AI for reading research papers: where you actually save hours
Not every paper deserves a full read, and AI is excellent for sorting before you sink an hour into text that turns out to be irrelevant.
A three-pass filter that works. First pass: the abstract and conclusion. Ask AI to summarize just these two sections and rate relevance to your question on a 1-5 scale. Papers scoring 1-2 go straight to the archive, no further work.
Second pass: methodology. For what's left, ask for a breakdown of exactly how the result was obtained: sample, method, limitations. AI is especially useful here for papers outside your main specialty, where the jargon of an unfamiliar method is hard to parse on your own.
Third pass: full reading with an AI assistant at hand. Only for papers that survive the first two filters, you read the text yourself, with a chat open for clarifying questions as you go. Not instead of reading, but alongside it: a question like "what does this term mean in the context of this specific paper" shouldn't have to interrupt reading with a trip to a separate search tab.
In practice, this cuts the time spent on a paper that turns out irrelevant from 40-50 minutes down to 5, and frees up hours for the small amount of material that's actually worth reading carefully, instead of spreading the same amount of time evenly across everything regardless of importance.
Worth a separate note on papers not in your native language: AI removes a barrier that used to cut off entire swaths of relevant literature simply because it was written in Japanese or German. Translation plus the same three-pass filter works the same way it does for text in your own language, though it's worth double-checking the nuances of phrasing in translation before quoting directly.
Prompts for document analysis
A library of working prompts for the extraction stage. Keep these on hand so you're not rewriting them every time you sit down with a new batch of sources.
Extract from this document:
1. The author's main claim in one sentence.
2. Three arguments they use to support it.
3. Any limitations or caveats the author acknowledges themselves.
4. One question the document leaves open.
Don't add anything of your own. Only what's explicitly in the text.
Compare these two documents: [document 1], [document 2].
Where do they agree, where do they directly contradict each other, and
where do they simply cover different aspects of the same topic without
overlapping? Don't smooth over contradictions. If they exist, name
them directly.
Here's my working hypothesis: [hypothesis].
Here's a source: [source text].
Does this source confirm the hypothesis, refute it, or is it unrelated
to it? Justify briefly. If the source partly confirms and partly
refutes it, say so, don't pick one side.
That last line in each prompt isn't an afterthought. A model defaults toward finding agreement and synthesis where an open contradiction between sources actually exists, and contradictions like that are usually the most interesting part for your own analysis. The third prompt is especially useful early on: it quickly shows whether a hypothesis is even worth further checking, or whether the material doesn't support it from the start.
What other AI research tools are worth knowing beyond NotebookLM
NotebookLM handles work with sources you've already gathered, but it doesn't help you find them. There's a separate class of tools built specifically for searching and doing a first-pass evaluation of research literature.
| Tool | What it does | When it's useful |
|---|---|---|
| Elicit | Searches papers by question, automatically extracts method, sample, and result into a table | A systematic review with hundreds of candidates, where you need a structured table rather than a list |
| Consensus | Searches for scientific consensus on a specific claim, shows the percentage of papers that agree or disagree | Quickly checking "is there even a scientific consensus on X" before going deeper |
| Semantic Scholar | Search with AI summaries and a citation graph between papers | Understanding who cites whom and where the real center of the discussion on a topic sits |
| Perplexity (Academic mode) | Open search with real-time cited sources | Quickly getting oriented in a topic that's new to you before you start gathering sources deliberately |
Practical advice: don't try to use one tool for every case. Each has its own role: Elicit and Semantic Scholar for search and initial sorting, NotebookLM for deep work with an already-selected corpus, plain ChatGPT or Claude for open-ended questions and explanations outside the scope of specific sources. Trying to cover three different jobs with one tool usually gives you a mediocre result at all three.
The order you use them in matters too, not just the set of tools itself. It's worth starting with a broad search through Elicit or Consensus to get the outline of the topic and not miss an obvious body of literature. Then narrow down to the genuinely relevant sources and only upload those to NotebookLM for deep work. The reverse order, trying to dump everything into NotebookLM upfront with no filtering, overloads the tool with low-relevance sources and lowers the quality of answers to the questions that actually matter.
Most of these tools are free at a basic level, and that's enough for a mid-sized project. Paying only makes sense once you hit a specific limitation: a cap on the number of documents you can upload, processing speed for a large corpus, no export in the format you need. Starting straight off with a paid subscription for a task you haven't even tried solving on the free tier usually means paying for features you'll never end up needing.
AI + Zotero and Mendeley: automating your bibliography
If you're already using Zotero or Mendeley to manage citations, there's no need to switch to something else for AI features: AI plugs into that process through plugins rather than replacing the tools themselves. The ecosystem around both services is mature enough that there's no reason to reinvent bibliography management from scratch. Current plugins can:
- automatically build a citation entry from an uploaded PDF, including cases with incomplete metadata;
- suggest tags based on a paper's content rather than just its title (the same caution principle as with personal notes: a closed tag list works more reliably than free generation);
- find duplicate entries that formally look different: a different author-name format, different capitalization in the title, a stray space.
The savings here aren't revolutionary, but they add up. On a single paper, automated bibliography management saves a couple of minutes. On a thesis with 200+ sources, that's hours of pure busywork that would otherwise go into formatting for APA, MLA, or Chicago style instead of into the substance of the work.
Taking notes while reading with AI
Notes taken while reading a source solve the same problem as notes in a personal knowledge base, just with a stricter accuracy requirement: a quote needs to be a quote, not a paraphrase passed off as one.
The practice that keeps this boundary intact: when extracting a claim, AI immediately flags whether it's a direct quote with an exact page reference or your own paraphrase. Blending these two types of notes is the source of most academic-honesty problems, which usually come from simply losing track of a phrase's origin a couple of weeks into the work, not from bad intent: you just can't remember anymore whether it was the author's exact wording or your own summary.
From there, these notes live exactly like any other atomic note: one thought, one entry, an explicit connection to other notes, rather than just getting filed away by source folder. AI Second Brain: Building a Personal Knowledge System covers the mechanics of this in detail. Here it's worth stressing just one thing: notes on sources shouldn't live in a system isolated from the rest of your knowledge base. A good idea from an article and a good idea from a conversation with a colleague should be able to meet in the same network, rather than living in parallel worlds of "work notes" and "thesis notes."
Synthesizing information from multiple sources: a step-by-step method
Synthesis is the hardest and most valuable part of working with sources, and it's the part that most often fails when AI is used carelessly. Asking a model to "draw conclusions from all these articles" almost guarantees vague text about nothing, because the model doesn't understand why you need this synthesis in the first place or what you're going to do with it next.
A method that actually works:
- State the synthesis question ahead of time, before sitting down with your sources. Not "what do the sources say about X," but "is there enough evidence to claim Y, and if not, what specifically is missing."
- Give AI all the extracted claims at once, not one source at a time. Synthesis requires seeing all the positions simultaneously, not sequentially, the same way you can't understand a map by looking at it one square inch at a time.
- Ask it to explicitly mark conflicts and gaps, not just agreements. "Where do the sources contradict each other, and can you guess why" is a question that usually surfaces the most interesting part of the material.
- Write the conclusion yourself, using the model's markup as a map, not as finished text. Same principle as with turning personal notes into an essay: the effort of formulating it has to stay on your side, or the synthesis ends up belonging to the model, not to you, and it'll show in the writing.
Let's walk through a hypothetical example. Say the synthesis question is: does microlearning in 15-minute daily sessions work better than one long weekly session, given the same total time. Five sources: two say yes, short intervals are more effective, one claims there's no difference for procedural skills, another says it all depends on the difficulty of the material, and the last one is actually about session frequency rather than length, a related but different question. A bad synthesis averages all of this down to "opinions differ." A good synthesis says: for declarative knowledge, the evidence favors short sessions; for procedural skills, the evidence is insufficient; and the literature actually conflates the question of frequency with the question of duration, which is an open methodological problem, not just researchers disagreeing.
The difference between a good synthesis and a bad one almost always comes down to whether the author named the contradictions openly or smoothed them over into a generic "researchers disagree." That's a sentence that says nothing and fits any topic under the sun, from learning effectiveness to the physics of black holes.
How to check AI summaries and avoid getting hallucinated at
In research work, the cost of a mistake is higher than in learning: a wrong fact that slips into a thesis, an article, or a work report costs you your credibility, not just a bad grade on an exam.
Three levels of verification, worth applying with increasing strictness depending on how important the fact is:
- A quick check. Ask the model to point to the exact location in the source: the page, the paragraph, where the fact came from. If it can't, that's already a warning sign, even if the fact itself sounds plausible.
- Manual re-verification. For facts that become the foundation of an argument rather than just background, open the original and check it yourself. Not occasionally "when it feels important," but systematically for everything that makes it into the final text.
- An independent second source. For the most important claims, find confirmation outside your original set of documents, ideally through a different method: not another AI summary of the same text, but a direct check against the original or an alternative source.
Let me tell a specific story that taught me not to cut corners on this step. I was gathering sources on the effectiveness of spaced repetition and asked AI to summarize data from half a dozen papers. The summary included a specific number: "studies show a 40% improvement in retention." The number sounded plausible and fit neatly into the piece I was writing at the time, so I almost left it as is. I went to check the source, and the original paper actually said "up to 40% under certain experimental conditions, with an average improvement of around 15% across the sample as a whole." The model didn't lie outright, it just took the most impressive number from the text and presented it as the general result. Since then I check by hand any number that makes it into a piece of writing from an AI summary, no exceptions, no matter how small or obvious it seems.
Worth a separate note on fact-checking numbers specifically: models are especially unreliable with exact numbers, dates, percentages, and sample sizes. The reasoning around a number can be flawless while the number itself is made up or pulled out of context, because for a language model, these are two different types of information with different reliability in reproduction.
How a student should organize a term paper or thesis with AI
Academic writing is a specific case of source synthesis with an added requirement of originality, and it's the easiest place to slide into the same full-delegation trap I wrote about in the piece on learning with AI.
The working boundary: AI helps at the stages of organizing sources, extracting claims, and checking logic and structure, but it doesn't write the substantive text in your place. In practice, that looks like:
- Literature review. AI helps find and structure sources (covered above); you write the review itself.
- Methodology. AI is good for checking that a method is described fully and reproducibly, bad for inventing the method for you.
- Data analysis. AI speeds up routine calculations and visualization; you do the interpretation of the results, because interpretation is exactly your contribution to the work.
- Final text. AI is good for structural critique of a draft, questions like "where does the argument break down," not for generating text from scratch or rewriting it in your place.
A thesis advisor reading text you didn't write will notice faster than you'd think. Not from the style (AI style isn't hard to fake at this point), but from the absence of exactly those small rough edges of understanding that only show up in someone who's genuinely gone through the material. At your defense, it surfaces even faster: one follow-up question about a methodology you formally described but didn't actually think through yourself is enough.
AI tools for analysts and consultants
Everything covered above applies beyond academia. An analyst preparing a market report from fifty sources and a consultant synthesizing client interviews into recommendations are structurally solving the same problem: scattered sources turn into verified claims, claims get synthesized into an argued conclusion.
The difference is mostly in the format of the sources (interviews, internal documents, competitor reports instead of research papers) and in how fast a result is needed: often days, not months. That makes systematic organization (a source registry, explicit flagging of verified versus unverified) even more important than in academic work. There's less time to notice and fix a mistake before it ends up in a client presentation, where the cost of a mistake is lost trust, not just a bad grade.
Another specific to consulting: sources are often confidential (a client's internal documents, interview recordings), which brings back the privacy question from the piece on second brains. Before uploading material like that to a cloud AI service, it's worth explicitly checking that specific tool's confidentiality terms and your company's policy on handling other people's data in AI services. That's something that can cost you a contract, not just a formal technical detail.
When a whole team is working through sources, not just one person
Everything described above is built around one person, but research projects at companies and academic labs are often run by a team of several people, each working through their own slice of the sources. This adds a new problem: the registry and notes need to be understandable not just to the author, but to colleagues who haven't read the original.
A practical rule for team work: every extracted claim in the shared registry needs to include not just the fact itself, but an explicit confidence rating from whoever extracted it. It's one thing for the author to write "confident, checked against the original." It's a completely different thing when they honestly admit: "AI produced this, didn't have time to check it myself." A colleague who's going to build on that claim in their own section needs to see the difference at a glance, not find out about it in a side conversation after the text is already written.
Second point: it's worth agreeing ahead of time on one shared prompt format for extracting claims (see the section above), rather than letting everyone invent their own. Otherwise one person's claims come out in one format, another's in a different one, and synthesis at the end turns into a separate task of forcing everything into a common shape, instead of gathering material in a consistent structure from the start.
Where everything you find should end up: a personal research base
Claims extracted from sources and notes taken while reading are exactly the material that should flow into a permanent knowledge base, rather than living only within the scope of one project and disappearing once the work is turned in.
The practical reason: the next research project almost never starts from absolute zero. A term paper's topic overlaps with a thesis topic, a report for one client echoes a pattern from a report for another. If sources are organized as one continuous, compounding system rather than a folder called "thesis_2026" you'll forget the day after your defense, these overlaps start working for you automatically, rather than only when you happen to remember you've read about this before.
AI Second Brain: Building a Personal Knowledge System covers how to build a system like that in general: atomic notes, explicit connections, AI search by meaning instead of by keyword. As applied to sources, the one addition: the registry from the organization section above should gradually flow into that same system, rather than living forever as a separate file.
How much time to budget for each stage
One reason work with sources stretches out over weeks with no visible progress: all the stages of the workflow get blended into one shapeless task of "dealing with the material." Splitting by time works better than splitting only by logic.
A rough breakdown for a mid-sized project (a term paper, a work report, an article): capture and initial sorting of sources takes a quarter of the total time, extraction and verification take half, synthesis and writing take the remaining quarter. It sounds counterintuitive that synthesis and writing take less time than extraction and verification, but in practice that's exactly how it goes: if claims are extracted carefully and verified ahead of time, writing moves fast, because you're not simultaneously formulating a thought and having to go re-check facts.
A common mistake in time allocation: spending disproportionately much on capture (an endless search for "just one more article on the topic") and too little on synthesis, which gets pushed to the last night before the deadline. If you notice you've been gathering sources for a week and haven't extracted a single claim yet, that's a signal to stop and start extracting from what you already have, rather than continuing to gather more.
A personal case: how I use AI in my own research work
While gathering material for this site and for Cruxly, I accumulated around a hundred sources: articles on Zettelkasten and second brains, research on spaced repetition, interviews with developers of similar AI tools. Without organizing them, I would've just drowned. Some sources I honestly read in full, some I ran through that same three-pass filter described above, and about forty percent got screened out on the abstract alone.
Let me single out where AI genuinely saved time. I had five long podcasts discussing PKM methodology, about an hour and a half each. In the past I would've either listened to all of them start to finish, losing evening after evening, or skipped them entirely, which is also a bad option, since some genuinely valuable information was in there. With transcription and AI summarization, exactly the scenario I built Cruxly for in the first place, I got through all five in one evening: a summary of each first, then a full listen only for the two that turned out genuinely relevant based on the summary. The remaining three went into the archive with a short claim attached, no regret about not listening to the rest.
This is literally why I started building that product in the first place. I felt the lack of a tool like this before I started building it. Not an abstract "the market needs this," but a concrete irritation at losing an evening to one podcast that had maybe five minutes of content I actually needed.
Common mistakes when working with sources using AI
- Asking for synthesis with no clear question. "Draw conclusions from these articles" with no statement of why you need those conclusions almost always produces vague text that fits any topic and therefore fits none.
- Not distinguishing a quote from a paraphrase. A couple of weeks later, there's no way to remember whether a phrase was verbatim or in your own words. That's not a minor detail, it's a direct risk of accidental plagiarism in academic work.
- Trusting numbers without checking them. A model's reasoning is usually more reliable than a specific number sitting inside it. Check numbers separately from the surrounding argument, always, even when they seem obviously correct.
- Using one tool for every task. NotebookLM doesn't search for sources, Elicit doesn't replace deep work with a corpus you've already gathered. Trying to cover the whole workflow with one tool usually gives a mediocre result at every stage.
- Skipping over disagreement between sources. A model defaults toward smoothing contradictions into a neutral "opinions differ." The contradiction between sources is almost always where the most interesting part of the analysis lives, and losing it means losing the main point.
- Storing material separately from your permanent knowledge base. A folder called "thesis_2026" disappears from view the day after your defense, along with everything accumulated in it. The next project has to find the same sources all over again.
- Uploading confidential material without checking a service's terms. Especially relevant for consulting and working with other people's documents: not every AI tool handles data that isn't yours equally responsibly.
Where to go from here
Working with sources is a distinct discipline layered on top of learning in general. If you haven't yet read the general breakdown of AI learning techniques (tutor mode, the Feynman technique, spaced repetition), start with How to Learn with AI: The Complete Guide.
If, after synthesizing your sources, you need to turn the extracted claims into a continuously growing knowledge base rather than a one-off document, AI Second Brain: Building a Personal Knowledge System covers the mechanics of storing and connecting them.
If your specific task is processing a lot of video or PDFs at once rather than just text articles, AI Video Note-Taking and AI Chat with PDFs cover those two cases separately and in detail. If you're deciding between specific tools for this, there's a comparison in the piece on NotebookLM and its alternatives.
Takeaways
- Working with sources isn't reading, it's a distinct process: capture, extract, verify, connect, synthesize. A break at any of these steps loses information.
- AI speeds up each step differently: a literature review through grouping positions rather than summarizing; reading through a three-pass filter instead of reading everything in full; synthesis through explicitly marking contradictions instead of smoothing them over.
- The cost of a mistake in research work is higher than in learning, so fact-checking isn't an optional step, it's a systematic habit applied with increasing strictness for more important claims. Especially for specific numbers: they lie more often than the reasoning around them.
- Claims and notes from your sources should flow into a permanent knowledge base, not disappear along with a closed project. The next research effort almost never starts from zero.
- There's no single tool "for everything," and it's not worth looking for one. NotebookLM, Elicit, ChatGPT, and a plain text-file registry each solve a different part of the workflow, and trying to reduce it all to one service usually means some stage ends up covered worse than the rest.
If, after six months of regular practice, synthesizing sources still takes you just as long as it used to, the problem is probably one of the skipped steps from the mistakes list above, not the tools themselves. It's worth going back and checking each step individually, rather than swapping one AI model for another hoping the new one compensates for a process that was never structured right to begin with.
Comments
No comments yet. Be the first.