Article

A Week of Co-Writing in Arabic With AI: Where the Machine Helped, Failed, and Surprised Us

By Khaled Editor • 2026-05-28 17:38

For seven days, we used an AI writing assistant to co-write in Arabic across a small but varied set of tasks: a 700-word personal essay built from voice-note transcripts, two short fiction scenes, a newsletter intro in Modern Standard Arabic, social captions, and a batch of headline ideas. We wrote in formal Arabic, Egyptian colloquial, and a lighter pan-Arab register. We logged what the system produced, what we kept, and how much editing each piece needed.

That matters because Arabic AI writing is still discussed more than it is tested in real editorial conditions. The central tension became clear almost immediately. The tool could generate smooth Arabic very quickly. But could it help without flattening voice, dialect, humor, and cultural detail into something generic? After a week, the answer was mixed. It sped up the edges of writing. It was far less reliable at the part readers actually remember.

How we ran the test

We kept the setup simple. One editor drafted prompts and one editor reviewed outputs. Over the week, we ran 48 prompts through one mainstream AI chatbot. We judged the results on four things: clarity, tone, originality, and edit time.

  • 11 outputs were usable with light edits
  • 17 outputs were salvageable only after heavy rewriting
  • 20 outputs were discarded

This was not a scientific benchmark. Different tools, different prompts, and different writers would produce different results. But it was enough to show where the system was genuinely useful and where it became expensive in a different way: not in money, but in cleanup time and lost voice.

Where it helped most: structure, options, and momentum

The best use case was not “write this for me.” It was “help me move.” When we gave the system rough material, it often helped us organize it faster than we would have on our own.

One example came from a voice note about a family iftar: a hot apartment, an absent uncle, a discussion about leaving the country, and a small moment of silence after dinner. The first AI draft was too stiff to publish. But the structure was useful. It split the material into three beats: the room, the argument, and the quiet after. That gave us a workable shape in minutes.

It was also good at generating alternatives. Headlines were a clear win. Out of 15 headline suggestions for one newsletter piece, three were strong enough to make the final shortlist. That may not sound impressive, but it cut the process roughly in half. The same was true for intros. If we already knew the argument, asking for five openings in different tones was often more useful than staring at a blank page.

Another practical strength was compression. When we pasted a dense paragraph and asked for a version that was 25 percent shorter without changing the meaning, the system usually did a decent first pass. It removed repeated qualifiers, tightened transitions, and exposed where we had buried the point.

“Useful for scaffolding. Not usable as voice.”

That note from our editing log sums up the strongest part of the week. The AI assistant worked best as a structuring tool, a variation engine, and a first-round editor. It was much weaker as a creator of finished Arabic prose.

Where it failed: dialogue, texture, and emotional truth

The weakest outputs appeared when the writing needed to sound lived-in rather than merely correct.

Dialogue was the clearest example. We asked for a short exchange between two sisters in Cairo arguing about rent. The result was readable, but unnatural:

“إنتِ تعرفين يا أختي أن الظروف صعبة، ولكن يجب أن نواجه الواقع.”

Every word is understandable. Very few people would actually say it that way in that scene. The sentence carries the grammar of written Arabic and the emotional temperature of a formal meeting. That mismatch came up often. When we asked for colloquial dialogue, the model drifted back toward formal phrasing. When we pushed harder for slang, it sometimes overdid it and produced speech that felt copied from the internet rather than overheard in life.

Reflective prose had a different problem. It defaulted to polished emptiness. We repeatedly got openings like this:

“في عالم يتسارع فيه كل شيء، تبقى الكتابة مساحة للتأمل والتواصل.”

The sentence is grammatical. It also says almost nothing. Arabic media already has many stock phrases. AI can reproduce them very fast. That is not a small weakness. It creates text that looks finished and feels hollow.

The system also struggled with specificity. Ask it for grief, and it reached for memory, warmth, and silence. Ask it for nostalgia, and out came coffee, old houses, and fading voices. Those images are not wrong. They are just interchangeable. A human writer usually earns feeling through detail. The model often skipped straight to the summary version of feeling.

Humor was another weak spot. Jokes came back explained. Sarcasm often became plain statement. A line meant to sound dry and sharp returned padded with extra words, as if clarity required removing all risk.

Arabic made the weaknesses easier to see

Some of these problems exist in English too. But Arabic makes them more visible because register carries so much social meaning.

Arabic is not one flat writing system. A single piece may move between Modern Standard Arabic, which is the formal written register, and a local spoken dialect. It may use a Quranic echo, a regional expression, an English loanword, or a phrase that only sounds natural in one city. Human writers navigate these layers almost automatically. AI often averages them.

The result is a kind of middle Arabic: clear, grammatical, and hard to place. It does not belong strongly enough to a person, a street, or an age group. For explanatory copy, that can be acceptable. For creative work, it quickly becomes a problem.

We saw this in small but telling choices. The system mixed formal verbs into casual dialogue. It used one region’s phrasing inside another region’s rhythm. When asked for youth-friendly Arabic with a little code-switching, it either turned too formal or leaned too heavily on English. It could produce correct sentences that were socially wrong.

“Readable, but no one I know would actually send this.”

That comment came up more than once in our notes. It points to an important distinction: fluency is not the same as credibility.

It is hard to say from one week whether these weaknesses came mostly from limited dialect training, from our prompts, or from both. That part remains uncertain. But the pattern was consistent. The farther the task moved from neutral exposition and toward voice, place, or intimate tone, the faster the quality dropped.

What surprised us

The biggest surprise was not in raw drafting. It was in revision.

When we pasted our own Arabic and gave very narrow instructions, the system became more useful. “Cut 20 percent.” “Give me three endings, one warmer, one sharper, one more direct.” “Mark repeated ideas.” “Rewrite this in simpler Arabic for a broad regional audience.” These prompts often produced something worth working with.

In other words, the tool was better at operating on existing material than inventing good material from nothing. That matters for writers who already have a voice and do not want to lose it.

It also did better when we named the audience precisely. “Write in clear Modern Standard Arabic for readers aged 20 to 35 across the Arab world” produced better results than “write naturally in Arabic.” “Give me spoken Egyptian but avoid exaggerated slang” was better than “make it colloquial.” Precision did not solve the problem, but it improved the hit rate.

One more surprise: the system was good at exposing our own habits. When asked to identify repeated rhetorical moves in a draft, it correctly flagged our overuse of throat-clearing phrases and long explanatory transitions. That does not make it a critic in the human sense. But it can be a practical mirror.

The real risk is not just error. It is sameness.

Obvious mistakes are easy to fix. Generic competence is harder to resist.

The danger in AI-assisted Arabic writing is not only factual error or awkward grammar. It is the spread of polished, anonymous prose. If many writers use the same systems for openings, transitions, summaries, and emotional framing, published writing starts to narrow. It becomes smoother and less distinctive. For Arabic, where public writing already carries pressure toward standardization, that risk is serious.

There is also an authorship issue. The more a writer depends on generated phrasing, the easier it becomes to outsource taste without noticing it. That is especially risky in creative work, where rhythm, silence, and local choice matter as much as information.

The promise is real too. AI can widen access to drafting support. It can help bilingual writers move between Arabic and English notes. It can help students compare registers. It can help editors simplify dense copy for readers who want clarity, not ornament. But these gains are strongest when the human writer already knows what the piece is trying to do.

What actually worked for us

  • Start with your own material. Notes, scenes, fragments, or a rough paragraph gave better results than empty prompts.
  • Name the register. Say whether you want Modern Standard Arabic, Egyptian colloquial, or a broad pan-Arab style.
  • Ask for options, not a final answer. Three openings are more useful than one “perfect” paragraph.
  • Use it for compression and restructuring. These were the most reliable tasks in our test.
  • Do not trust dialogue on first pass. Check it with a real speaker from the relevant place.
  • Read everything aloud. Arabic rhythm exposes weak prose quickly.
  • Avoid imitation prompts. Asking a model to mimic living writers is a creative shortcut and an ethical one too.

After one week, the boundary looked clearer

Our experiment did not end with fear of replacement or with easy praise. It ended with a narrower, more useful view. In Arabic, AI is already helpful for setup work: structuring notes, testing tone, shortening drafts, and generating alternatives. It is still unreliable at the center of creative writing, where voice, dialect, and cultural instinct do the real work.

The practical lesson is simple. Use AI to widen your options. Do not let it decide your sentences for you. In Arabic especially, the writer’s ear is still the difference between text that merely reads well and text that feels true.