Should Researchers Share Unpublished Work With AI? Only With Real Safeguards
A recent online discussion asked whether researchers should trust commercial AI tools such as ChatGPT with unpublished math. It was not a formal study, but it captured a real change in academic life. These systems are now useful enough that people want to use them on work that is still private, incomplete, and potentially career-defining.
This matters far beyond mathematics. The same choice faces PhD students with draft chapters, academics with preprint-ready papers, writers with unfinished books, and founders with early inventions. The tension is simple: AI can provide fast feedback and save time, but once original work enters a third-party system, privacy, control, and credit become harder to manage.
My view: do not share unpublished work by default
Researchers should not paste unpublished work into general consumer AI tools by default. They should use AI on private work only when the protections are clear, the exposure is limited, and the benefit is worth the risk.
That position is not anti-AI. It is basic professional caution. Researchers already treat lab data, grant drafts, peer review, and collaboration notes as sensitive. An unpublished proof, method, or argument deserves the same discipline.
Trust should be technical, not emotional
When people say they “trust AI,” they usually mean something narrower and more practical. Who can access the input? How long is it stored? Is it used for training? Can administrators or reviewers see it? What happens if the terms change or the provider has a security incident?
The facts: protections vary widely by product, plan, and institution. Some enterprise or API tools offer stronger limits on retention and training use than consumer chatbots. The interpretation: most ordinary users do not have a clear view of those differences. My opinion: if you cannot explain the data path, you should assume the risk is higher than it looks.
Why unpublished work is different
Unpublished work is not just another draft. It often contains the part that creates value: the original proof, the new result, the fresh framing, the patentable claim, the argument that establishes priority. Once that material leaves your own files and enters an outside system, you lose some control over who may see it and how it may be handled.
In pure mathematics, the main risk may be less about personal privacy and more about priority and timing. In medicine, biotech, or engineering, confidentiality can also affect compliance, patents, and contracts. Different fields face different dangers, but the common issue is the same: unpublished work has a longer tail of consequences than a routine prompt.
The promise is real
There is a reason researchers are tempted. AI tools can help rewrite dense prose, check notation, suggest examples, translate jargon into plain English, generate test code, or point to assumptions that need to be stated more clearly. For non-native English speakers and people without easy access to mentors, that can be genuinely useful.
Used carefully, AI can also help without exposing the core invention. A researcher can ask for LaTeX help, request a clearer explanation of a known theorem, or test a toy version of an argument with the novel part removed. That captures much of the convenience without giving away the most sensitive material.
But usefulness is not the same as reliability. These systems can still invent citations, misread a proof, or smooth over a real gap with confident-sounding text. Even in a fully private setup, the output still needs expert checking.
Credit may be the bigger risk
Privacy gets most of the attention. Credit may be the harder problem. In research, value depends not only on being first, but on being able to show what you contributed, when you contributed it, and under what conditions the work was developed.
If a researcher heavily uses AI on an unpublished manuscript or proof, several questions follow. What should be disclosed to coauthors, supervisors, editors, or patent counsel? How should the researcher document the timeline? If a later dispute arises, what records exist? Some journals, conferences, and universities now ask for disclosure of AI assistance, but the rules are still uneven.
AI does not need to “steal” an idea for credit to become blurry. Confused records, unclear authorship boundaries, and weak documentation can do the damage on their own.
For inventors and startup researchers, the credit question can spill into law. A quick prompt that includes the key inventive step may seem harmless, but later it can complicate confidentiality claims or inventorship discussions. That is a high price to pay for a few minutes of convenience.
The strongest counterargument
The strongest counterargument is practical and fair. If AI can act as a quick second reader, why should only well-connected researchers use human networks while everyone else avoids a cheap tool? That concern matters. A blanket ban would protect confidentiality, but it could also protect existing inequality.
There is also a simple reality: many researchers are already using these systems quietly. Refusing to discuss the practice will not stop it. It will only ensure that people use the tools without standards.
A better rule than “never”
The right rule is not “never use AI.” It is “share the minimum, and protect the core.” That means separating low-risk assistance from high-risk disclosure.
A simple example shows the difference. Asking a model to clean up LaTeX for a published proof is one thing. Uploading a full new proof and asking whether it is original is something else. The first uses AI as a tool. The second exposes the asset itself.
- Usually low risk: formatting help, grammar, presentation, public-source material, generic coding patterns, and questions that do not reveal the novel result.
- Sometimes acceptable with care: short sanitized excerpts, toy examples, or reformulated problems that remove the central insight and any sensitive data.
- Usually high risk: full unpublished manuscripts, proofs, confidential data sets, grant drafts, peer reviews, invention disclosures, patent claims, and collaboration notes pasted into consumer tools.
If the tool is locally run, institutionally approved, or covered by a contract with real privacy protections, the line can move. But that should be a conscious decision, not a habit driven by convenience.
Five questions to ask before you upload anything
- What is the worst-case cost if this leaves your control? Think about publication priority, collaborator trust, legal exposure, and reputation.
- Do you know the exact terms of the exact product you are using? Brand names are not enough. Plans and settings matter.
- Can you ask the question without the secret part? Often you can remove the novel lemma, undisclosed result, or identifiable data and still get useful help.
- Would you be comfortable disclosing this use later? If you had to explain it to a coauthor, editor, supervisor, or patent lawyer, would it sound sensible?
- Is there a safer alternative? A local model, a secured API workflow, an enterprise account, or a trusted human reader may solve the same problem with less risk.
Institutions need to do their part
Too much of this decision still sits on individual researchers. Many universities and labs talk about AI in broad terms but offer little concrete guidance on unpublished work. That is not good enough.
Institutions should publish short, usable rules: which tools are approved, which kinds of material are off-limits, when disclosure is required, how records should be kept, and who is responsible if a breach occurs. Researchers should not have to decode vendor policies on their own.
A sensible default
Researchers should treat unpublished work as confidential by default. Use AI for editing, explanation, and routine support when you can do so safely. Do not hand over the core of an original result unless the protections are specific, enforceable, and worth the trade.
The key point is simple. Trust is not a feeling about a tool. It is a judgment about contracts, storage, access, and consequences. If you would not share the material widely before publication, do not upload it casually just because the interface feels private.