金融翻訳者の日記/A Translator's Ledger

自営業者として独立して十数年の翻訳者が綴る日々の活動記録と雑感。

「こてこての関西弁」は英語で何と言うのか?(2026年8月4日)

「彼女はこてこての関西弁で次のように尋ねた」という、くだけた物語や随筆の文脈なら、第一候補は in full-on Kansai dialect です。

Then, in full-on Kansai dialect, she asked, "…"

または、

She asked in full-on Kansai dialect, "…"

なぜ full-on なのか

「こてこて」は、単に「強い」「濃い」というだけの表現ではありません。「こてこてのソース」「こてこての化粧」と同じように、薄められていない濃さ、いかにもそれらしい典型性、そして軽い可笑しみを含んでいます。「絵に描いたような関西弁」「これぞ関西弁」に近い語感です。

full-on は、程度が非常に強いことを表す口語的な形容詞です。Cambridge Dictionary の語釈は次の通りです。

full-on — very great or to the greatest degree
(非常に大きな程度で、あるいは最大限に)
Cambridge Dictionary, "full-on"

そのため、「こてこて」が持つ濃さ、典型性、口語的な勢い、軽い可笑しみを比較的よく表せます。

ただし、full-on Kansai dialect 自体を英語の決まり文句として扱っているわけではありません。一般的な口語表現である full-on を使い、「こてこて」の調子を英語として自然に再現したものです。

また、full-on を使っても、話し手が意識的に関西弁を誇張して演じていた、という意味にはなりません。あくまで語り手から見て、関西弁の特徴が前面に出ていたことを表します。

なぜ accent ではなく dialect なのか

この文では、直後に実際の発話が続きます。そこで問題になるのは発音だけではありません。関西特有の語彙、語尾、文法、言い回しまで含まれます。

Oxford Advanced Learner's Dictionary の dialect の語釈は次の通りです。

dialect — the form of a language that is spoken in one area with grammar, words and pronunciation that may be different from other forms of the same language
(ある地域で話される言語の形態で、文法・語彙・発音が、同じ言語の他の形態とは異なりうるもの)
Oxford Advanced Learner's Dictionary, "dialect"

文法・語彙・発音のすべてを含む、と明記されています。これに対して accent は主として発音上の特徴を指すため、この場合は dialect のほうが適切です。

なぜこの修飾句が重要なのか

台詞を英語で示す以上、関西弁特有の語尾や言い回しは、通常、そのまま英語には現れません。したがって、「こてこて」という濃さを読者に伝える役割は、主として in full-on Kansai dialect という修飾句が担います。

中立的な strong でも基本的な意味は伝わりますが、「こてこて」の口語的な勢いや軽い可笑しみは弱くなります。

文体による代替案

中立的で説明的な文章

in a strong Kansai dialect

明快で安全な表現です。ただし、「強い関西弁」という中立的な意味に近く、「こてこて」の軽い可笑しみは薄れます。

発音上の訛りだけを強調する場合

in a thick Kansai accent

thick は、強く目立つ訛りについて使われます。ただし、accent が表すのは主として音声上の特徴です。関西特有の語彙、語尾、文法、言い回しまで含めたい場合には、意味の範囲が狭すぎます。

broad を使う場合

in broad Kansai dialect

これも英語として成立します。broad には、方言や訛りについて、地域的特徴が強く、はっきり表れていることを示す用法があるからです。

Cambridge Dictionary は broad をこう説明しています。

If someone has a broad accent, it is strong and noticeable, showing where they come from.
(ある人が broad accent で話すという場合、その訛りは強く目立ち、どこの出身かが分かるものである)

例文:He spoke with a broad Australian accent.
(彼は強いオーストラリア訛りで話した)
Cambridge Dictionary, "broad"

そして Oxford Advanced Learner's Dictionary は、dialect の項の用例に次を挙げています。

She spoke in broad Yorkshire dialect.
(彼女は濃いヨークシャー方言で話した)
Oxford Advanced Learner's Dictionary, "dialect"

これは無冠詞で dialect を修飾する形です。したがって in broad Kansai dialect は、語法として辞書の用例と同型であり、まったく問題ありません。

それでもこの文脈で第一候補にしないのは、語法の正しさとは別の二つの理由によります。

一つは、この broad の語義が、一般的な「広い」という意味からは推測しにくいことです。Cambridge の語釈が accent を軸に立てられているように、この用法は accent と結びついた形のほうが理解されやすく、読者によっては即座に意味を取れない可能性があります。

もう一つは、broad が地域色の強さは表せても、「こてこて」が持つ口語的な勢いや軽い可笑しみまでは十分に拾わないことです。

実務上の注意

読者が関西を知らない可能性があるなら、台詞を導くこの一文で初めて地理を説明するのではなく、それ以前の初出箇所で一度説明しておくのが最善です。たとえば、

Kansai, a region in western Japan that includes cities such as Osaka and Kyoto

などと説明しておけば、台詞を導く一文は、

Then, in full-on Kansai dialect, she asked, "…"

と簡潔に書けます。

どうしてもこの位置で説明しなければならない場合は、ダッシュによる挿入を使います。

Then, in full-on Kansai dialect—the way people speak in the Kansai region of western Japan—she asked, "…"

ただし、これは台詞の直前に置く説明としてはやや長いため、できれば前の箇所で説明を済ませるほうがよいでしょう。

また、直後の台詞を米国南部方言やコックニーなど、別の英語方言に置き換えるのは避けたほうが安全です。その土地に固有の地域的・社会的含意まで持ち込んでしまうからです。通常は台詞を標準的な英語で示し、地の文の in full-on Kansai dialect で関西弁らしさを伝えます。

結論

「彼女はこてこての関西弁で次のように尋ねた」という、くだけた物語や随筆の文脈での第一候補は、

Then, in full-on Kansai dialect, she asked, "…"

です。

中立的な文章なら in a strong Kansai dialect、発音上の訛りだけを問題にするなら in a thick Kansai accent が適切です。in broad Kansai dialect も辞書の用例と同型で、語法上の問題はありません。

しかし、「こてこて」が持つ濃さ、典型性、口語的な勢い、軽い可笑しみまで表すなら、in full-on Kansai dialect がこの文脈では最も適切です。


Sources:

 

ご注意
この記事は、複数の生成AIとの対話を何度も往復しながらまとめたものです。
生成AIの回答には、誤答や情報の混乱、論理の不整合、事実誤認、情報の抜け落ちが含まれる可能性があります。
回答を鵜呑みにせず、あくまで考えるための材料としてご活用ください。

Still Our Teacher, Still Her Students—53 Years Later (August 2, 2026)

Yesterday there was a reunion with Junko T.—Ms. T. to us—who taught us English in seventh grade. Nine people attended, including her. For me it was the first time I had seen her in a full 53 years.

I had never studied a word of English in elementary school, so my first encounter with the language was, quite literally, Ms. T.’s class. The first sentences we learned were, I think, “This is a pen” and “Is this a pen?” She would read them aloud almost as if she were singing, with pronunciation and intonation so exaggerated you might have thought you were watching the Takarazuka Revue—Japan’s all-female musical theater company, famous for its lavish, highly stylized productions.

She got English into us first as sound and rhythm, and only then moved on to reading, writing, and speaking. Looking back, I suspect that was remarkably innovative for the time. Her classes were distinctive in another way, too: she taught in full-on Kansai dialect—the speech of the Kansai region in western Japan—even though our school was in Saitama, well outside Kansai. When I asked her about it yesterday, she told me she was from Kyoto.

I had also started listening to Kiso Eigo (“Basic English”), a radio course on NHK, Japan’s public broadcaster, but I rarely lasted more than a few days at a time. In the end, Ms. T.’s classes were the only real English instruction I had.

The middle school I attended was a large one. In eighth grade it was split in two, and I moved to the newly opened school while Ms. T. remained at the original one. Today she chairs the board of trustees at a private school in Saitama, and when I read her profile on the school’s website, I learned that before becoming an English teacher she had worked as an interpreter at the World Bank. Then it clicked: those unusual classes had not come about by accident.

The gathering was relaxed and friendly. Now 85, Ms. T. arrived almost exactly on time and opened with, “I have to go to the bathroom a lot, so let me sit over here.” She chose a seat on the side opposite the place of honor and sat down, straight-backed and alert. Still speaking in Kansai dialect, she launched in with something like, “Oh, K-kun! How’ve you been?” And just like that, we were back 53 years.

For most of us former students, by contrast, it was the first time we had seen one another since graduation. We were wearing name tags, but at first we could see only faint traces of the faces we remembered, and we began by addressing one another formally. Then someone brought out an old album, and as we talked our way through it, the formality gradually fell away. It did not take long for the polite -san after our names to give way to the familiar -kun of our school days.

When Ms. T. talked about running the school, she had no tales of hardship—only stories that clearly delighted her.

“Excellent teachers come to work for us after retiring from well-known prefectural high schools. Isn’t that wonderful?”

That was the general tone. Here was a board chair past eighty, still teaching middle and high school students alongside a native English-speaking instructor. We were the ones who came away energized.

And yet: “Still, I suppose I can only go on like this for another year or two,” she said—startlingly matter-of-fact about her own age.

A little past halfway through the gathering, Ms. T. said, “Just going to the bathroom,” and got up. A member of the restaurant staff came over to our table and whispered something to the organizer.

“Oh no—surely not.”

“What? What is it?”

It turned out that Ms. T. intended to pay for the entire meal.

“That’s completely backwards!”

The table was still in an uproar when she returned.

“Come on, Ms. T. You can’t do that.”

“Now listen,” she said. “I’m still working, and all of you are my students. Of course the teacher pays! But you’re all working adults, so paying nothing wouldn’t be right either. Make it 1,000 yen each.”

That was that. We ended up enjoying a multicourse French dinner for 1,000 yen apiece, plus another 500 yen each toward a gift for her.

“Let’s do this again next year. We can split the bill next time,” she said, and with that the gathering came to an end.

After seeing her off, we went on to a coffee shop, where someone said:

“When you get right down to it, in Ms. T.’s eyes, we’re still the middle schoolers we were back then.”

Oddly enough, it made perfect sense to all of us.

I’m already looking forward to next year.

https://tbest.hatenablog.com/entry/2026/08/02/123455

Looking for an excellent professional translator? Contact Tatsuya Suzuki at tbest.backup [at] gmail.com (Replace [at] with @)

いつまでたっても「先生」と「生徒」――53年ぶりの再会(2026年8月2日)

昨日は、中1の時に英語を教えてくださったT淳子先生を囲む会があった。先生を含め、集まったのは9人。僕が先生にお会いするのは、実に53年ぶりだった。

小学校時代に英語を習ったことは一切なかったので、僕と英語との出会いは、文字通りT先生の授業だった。This is a pen. Is this a pen? が最初に学んだ文だと思うが、これをT先生は宝塚歌劇とはかくやと思うような、かなり大げさな発音とイントネーションで、歌うように読み上げられた。英語をまず音とリズムとして体に入れ、そこから読み、書き、話すことへと広げていく。そんな教え方は、当時としては相当に画期的だったのではないかと思う。しかも授業はこてこての関西弁(僕が学んだ中学は埼玉県)という、非常に独特なものだった(昨日お聞きしたら京都のご出身だった)。僕も「基礎英語」を聴き始めていたものの、ほとんど三日坊主だったので、結局僕の英語学習はT先生の授業のみだったことになる。

僕がいた中学は大規模校で、中2の時に学校が分離し、僕は新しくできた学校へ。先生はそのまま同じ中学に残られた。今は埼玉のある私立学校の理事長をなさっていて、ウェブページで経歴を確認すると、英語教師になる前は世銀で通訳をなさっていたことがわかり、なるほど、あの独特な授業は偶然に生まれたものではなかったのだ、と腑に落ちた。

会は和やかに進んだ。85歳の先生はほぼ定刻にお見えになると、開口一番「あたし、トイレが近いからこちら側に座らせてね」と上座の反対側を選ばれ、シャキッと座られた。関西弁丸出しのまま「あら~○○君、どうやった~?」てな調子で、一気に53年前に逆戻りだ。

一方、我々元生徒のほとんどは、互いに顔を合わせるのが卒業以来初めてだった。名札をつけてはいたものの、最初は相手に昔の面影を感じる程度で、会話も敬語から始まった。誰かが持ち込んだアルバムを見ながら話すうちに次第に打ち解け、「さん」づけが「くん」づけに変わるまでに、さほど時間はかからなかった。

学校経営について、先生の口から出てくるのは苦労話ではなく、楽しそうな話ばかりだった。「有名な県立高校を定年退職しはった優秀な先生方が、来てくれはるのよ~。うれしいやないの」てな調子。80歳を過ぎた理事長が、いまだにネイティブ講師とペアで中高生向けの授業をしていると聞き、こちらがかえって元気をもらってしまった。その一方で、「でも、せいぜいあと1、2年だと思うけどね」と、自分の年齢については驚くほど冷静だった。

会が半ばを過ぎるころ、先生が「ちょっとトイレ」と言って立つと、レストランの方が我々の席に近づき幹事に何やら耳打ちする。「え~、そんなあ」「どうしたどうした?」と尋ねると、先生がこの会の食事代を全額お支払いになるという。「それは本末転倒だろ!」と騒然としているところに先生が戻られる。「そりゃないっすよ、先生」と言うと、T先生は、「い~い?あたしは今働いていて、皆さんはあたしの生徒なの。先生が払うのが当たり前じゃないの! といっても皆さん社会人だからゼロはないわよね。じゃあ、一人1000円ということにしなさい」とピシャリ。結局僕らは、1000円プラス先生へのお土産代500円でフレンチのフルコースをいただいたのであった。「来年もやろうやないの。今度は割り勘でいいからね」でお開きとなった。

先生をお送りしたあと、二次会で入った喫茶店で、誰かが言った。

「結局、T先生から見れば、僕らは今もあの頃の中学生なんだよね」

みんな、妙に納得した。

来年が楽しみだ。

tbest.hatenablog.com

Question Agreement, Use Disagreement: Cross-Checking with Generative AI — The Case of Translation(July 30, 2026)

Suppose you have in front of you a source text that needs checking, and a translation or summary made from it. Having an AI translate a document, or condense a long report: that is the kind of work I mean. It makes no difference whether the translation runs from Japanese into another language or the other way around. It could equally be a human translation you have asked an AI to review. Either way, everything to be verified is already there. The question is where to start checking.

If that is the situation, then the surest way to find and remove the errors hallucination causes — added information, misreadings, omissions and the like — is to go back to the source and check it with your own eyes. But that takes time and attention.

So the cost-effective second best, I would suggest, is this: put the same problem separately to several generative AIs you have reason to trust, and cross-check the answers.

That said, I do not think it is simply a matter of counting votes in favor. When two newspapers carry the same story off the same wire, that does not corroborate it. There is only one source behind them both. What agreement tells you depends on how far the two sources overlap. The same goes for people: when two people who studied under the same teacher agree, that carries less weight than when two people with entirely different backgrounds do.

With generative AI, this overlap of sources is large. That does not mean the models draw on a narrow range. If anything, the opposite is true. The major models are trained on vast bodies of data that include what is publicly available on the internet. No one outside can see how those datasets are composed, but as long as the same open web is one of their main sources, the wider each model reads, the more likely it is that two of them have read the same things. Two human beings, by contrast, draw on a range far narrower than an AI's, and one that differs from person to person, so their reading overlaps comparatively little. Breadth does not create independence.

Nor is it any guarantee that a model has read countless accounts that contradict each other. A model does not draw lots to decide which account to believe; it learns what sounds plausible. If a mistaken explanation of some point is widely circulated, several models may lean toward that same majority account. The voices of countless teachers collapse into a single teacher, weighted by frequency and reputation.

And unlike with a person, an AI's output gives you nothing to go on in deciding how far to discount it — nothing like knowing who someone studied under. In checking a translation, you cannot see from outside whether an answer came from reading the source properly or was pulled along by some stock misunderstanding the model had learned. Ask it for a citation on a matter of external fact and an answer does come back, but the citation itself may even be fabricated. That much, at least, is relatively easy to test: check not only whether the cited source exists, but whether what it says there actually supports the claim.

So adding more AIs to the cross-check does not necessarily improve the result in proportion. Ask similar questions of AIs that learned from similar material, and they may repeat the same error. What matters is not just how many AIs you use, but how far you spread the routes by which they reach an answer.

Start by giving several models the same source text, the same translation, and instructions identical in content, then have each check the whole of it independently, without seeing the others' answers. Choose models from different developers where you can. Only then can you compare where the answers agree and where they diverge. Then, to reduce what gets missed, send them through a second time with the angles of inspection divided: one looking for drift in meaning and logical relations, the other for numbers, proper names, negation, and omissions.

Leaving the source text alone and changing only the language you put your questions and instructions in can serve as a further measure. That alone does not make the answers independent, but it is one way to reduce the chance of their converging on the same error.

What you can draw from the answers differs completely depending on whether they diverge or agree. If they diverge, one of them may be wrong. Both may be wrong. The source itself may allow both readings. Whichever it is, you can say for certain that there is something there to be looked at.

The trouble is agreement. Since errors can correlate, agreement does not distinguish “both right” from “both carrying the same error.” How much corroboration you get from agreement depends, then, on how far the errors correlate. And how far they correlate is not something you can see from outside.

The point of cross-checking with several AIs, then, is not to set the answers side by side and survey them, but to put the places where they diverge back to both models and make them argue it out from the source. The moment a passage is presented as contested, what had been skimmed becomes an object of scrutiny. In practice, one model will sometimes give its reasons and revise its reading; sometimes both hold to different readings, and when a human finally works through the source, both readings turn out to stand. What is happening, presumably, is that one model's answer confronts the other with an objection it would not have arrived at on its own. At this point, rather than asking “which of you is right?”, it is better to ask “what would the source have had to say for the other reading to hold?” — because that asks the model to analyze the other reading rather than defend its own. This part is the human's job.

Agreement that both sides converge on after each has given reasons grounded in specific passages of the source can, I think, be trusted more than agreement that was there from the start. On one condition: that it is not the result of one side falling in behind the other without grounds. It is worth looking at whether the side that changed its mind can explain its own misreading with reference to the source. This, too, is the human's job.

Cross-checking among AIs, then, should in my view be treated not as a means of establishing what is correct, but as a means of putting the places a human needs to check in order of priority. Every point of divergence gets checked, without exception. Beyond that, even where the answers agree, the elements where an error costs most — numbers, proper names, negation, omissions — get spot-checked. Agreement does not rule out a shared error. Cross-checking is not a tool for skipping the source. It is a tool for deciding where to start checking.

https://tbest.hatenablog.com/entry/2026/07/30/174948

Looking for an excellent professional translator? Contact Tatsuya Suzuki at tbest.backup [at] gmail.com (Replace [at] with @)

一致を疑い、食い違いを生かす――生成AIによるクロスチェック(翻訳を事例として)

今、確認すべき原文と、それをもとに作られた訳文や要約が手元にある場面を想定する。文章をAIに訳させる、長い資料を要約させる、といった作業がこれに当たる。訳す方向が、日本語から他言語へであっても、他言語から日本語へであっても、話は変わらない。人間が訳したものをAIに点検させる場合でもよい。いずれにせよ、確かめるべきものは最初からそこにあり、問題は、どこから確認するかだ。

だとすれば、ハルシネーションによる情報の追加、読み違い、訳抜けなどの誤りを見つけ、取り除く最も確実な方法は、原文に戻り、自分の目で確認することだ。しかし、それには時間と注意力を要する。

したがって、費用対効果の高い次善策は、自分が信頼できる複数の生成AIに同じ問題を別々に検討させ、回答をクロスチェックすることではないか。

ただし、単純に賛成票を数えればよいわけではないと思う。同じ通信社の配信を引いた二つの新聞が同じことを書いていても、裏が取れたことにはならない。もとの出どころが一つだからだ。一致から何が言えるかは、両者の出どころがどれだけ重なるかで決まる。これは人間についても同じで、同じ師に学んだ二人の意見が揃っても、まったく別の経歴を持つ二人の意見が揃った場合ほどの重みはない。

生成AIの場合、この出どころの重なりが大きい。これは、参照する範囲が狭いという意味ではない。むしろ逆である。主要な生成AIは、公開インターネット上の情報を含む大規模なデータで学習している。学習データの詳しい構成は外から分からないが、同じ公開ウェブを主要な供給源の一つとしている以上、読む範囲が広がるほど、二者が読んだものは重なりやすくなる。一方、二人の人間が参照する範囲はAIよりもはるかに狭く、しかも人によって異なるため、比較的重なりにくい。網羅性は独立性を生まない。

矛盾しあう無数の記述を読んでいることも、担保にはならない。モデルはどの説を信じるかをくじで決めているのではなく、何が「もっともらしいか」を学んでいるからだ。ある論点について誤った説明が広く流通していれば、複数のモデルが同じ多数派の説明へ寄ることがある。無数の師の声は、頻度や評価に応じて重みづけされた一人の師へと合成されてしまうのだ。

しかも人間と違って、生成AIの出力からは、割り引き方を決める手がかり――あの人は誰に学んだか、というような情報――が得られない。訳文の点検であれば、ある答えが原文をきちんと読んで出てきたものか、学習した典型的な誤解に引きずられたものかが、外からは見えない。また、外部の事実について典拠を尋ねれば答えは返ってくるが、その典拠自体が捏造されている場合すらある。もっとも、これは比較的確かめやすい。示された典拠が実在するかだけでなく、そこで述べられている内容が、本当にその主張を支えているかを確認すればよいからだ。

だから、互いにチェックし合うAIの数を増やしても、効果がそれに比例して高まるとは限らない。似た情報から学んだAIに似た問い方をすれば、同じ誤りを繰り返す可能性があるからだ。重要なのは、AIの数だけでなく、答えに至る経路を分散させることである。

まず、複数のモデルに同じ原文と訳文を示し、同じ内容の指示を与えて、互いの回答を見せずに全文を独立して点検させる。モデルは、できるだけ開発元の異なるものを選ぶ。こうして初めて、回答の一致と食い違いを比較できる。そのうえで見落としを減らすため、一方には意味や論理関係のずれを、もう一方には数字、固有名詞、否定、訳抜けを重点的に調べさせるなど、点検の観点を分けて再検討させる。

さらに、原文はそのままに、AIへの質問や指示に使う言語だけを変えることも補助策になりうる。それだけで回答が独立するわけではないが、同じ誤りに揃う可能性を減らす一法ではある。

そして、複数のAIの回答が食い違った場合と一致した場合とでは、引き出せることがまるで違う。食い違ったなら、一方が誤っているのかもしれず、両方が誤っているのかもしれない。原文自体がどちらにも読めるのかもしれない。いずれにせよ、検討すべき何かがそこにあることは確実に言える。

問題は、複数のAIの回答が一致した場合だ。誤りが相関しうる以上、一致は「両方正しい」場合と「同じ誤りを共有している」場合とを区別しない。つまり、一致がどれだけの裏づけになるかは、誤りの相関の程度に大きく左右される。そしてその程度は、外からは分からない。

だから、複数のAIによるクロスチェックの要点は、答えを並べて眺めることではなく、食い違った箇所を双方のAIに突きつけ、根拠を挙げて論じさせることにある。争点として示された瞬間、それまで流し読みされていた箇所が精査の対象に変わる。実際、一方が理由を挙げて読みを改めることもあれば、双方が異なる読みを維持し、最終的に人間が原文を検討した結果、どちらの読みも成立すると分かることもある。一方の回答によって、他方は自力では思いつかなかった異論に直面するのだろう。このとき「どちらが正しいか」と問うより、「相手の読みが成立するには、原文がどうなっていればよかったか」と問うほうがよい。自説の防衛ではなく、相手の読みの分析を求めることになるからだ。ここが人間の仕事である。

双方が原文の具体的な箇所に基づいて理由を示し、そのうえで収束した一致は、最初からの一致より信用してよいと思う。ただし、一方が根拠なく他方に追随した結果ではないことが条件である。意見を変えた側が、自分の読み違いを原文に即して説明できるかどうかを見るとよい。これもまた、人間の仕事である。

このように、AI同士のクロスチェックは、正しさを確定する手段ではなく、人間が確認すべき箇所に順序をつける手段と位置づけるべきだと思う。食い違った箇所は必ず確認する。そのうえで、一致した箇所についても、数字、固有名詞、否定、訳抜けのように誤りの影響が大きい要素は、抜き取りで確かめる。一致は、共有された誤りを排除しないからだ。クロスチェックは、原文確認を省くための道具ではない。どこから確認するかを決めるための道具である。

tbest.hatenablog.com

 

 

Why It’s So Hard to Compare Notes on Generative AI(July 27, 2026)

Last week, on separate occasions, I happened to hear several people share their impressions of generative AI.

A: “ChatGPT gets things wrong a lot, so lately I’ve been using Gemini.”

B: “Claude isn’t very reliable, so I mostly use ChatGPT.”

C: “Hallucinations are a concern, but at least the numbers are right.”

I did chime in when C said that, speaking from my own experience.

“It can get numbers wrong too. Or rather, it can get the order of magnitude wrong or skip figures in a table. I don’t think you should trust it blindly.”

With A and B, however, I chose not to challenge either of them, even though my own impressions were different.

To begin with, I don’t know when either of them used it. That, I think, is the biggest issue. Generative AI is improving pretty quickly. Even “ChatGPT” covers a range of models, and I don’t know which one they used. Nor do I know what they were trying to get it to do or what kind of prompt produced the result they described.

Then there is a further question: how many times did it get something wrong without their noticing, leaving them to assume the answer was correct? There is no real way to ask them that, and they probably wouldn’t know themselves.

I have no reason to doubt that all three ran into the kinds of situations they described. But it is hard to conclude from those experiences which of ChatGPT, Gemini, and Claude is accurate and which is unreliable.

With most products and services, quality and characteristics are relatively stable. Different users can therefore identify much the same strengths and weaknesses and arrive at a broadly shared assessment.

Generative AI is different. Even within the same service, the underlying system varies depending on when you use it and which model is involved. The results also vary widely depending on what you ask it to do and how you instruct it. And even the user’s assessment of the AI itself changes depending on whether they can spot its mistakes.

That, I found myself thinking over drinks, is exactly what makes generative AI so hard to evaluate.

(Oh yes, all three conversations took place over drinks.)

https://tbest.hatenablog.com/entry/2026/07/27/161554

Looking for an excellent professional translator? Contact Tatsuya Suzuki at tbest.backup [at] gmail.com (Replace [at] with @)

共有しにくい生成AIの評価(2026年7月27日)

先週、たまたま別々の機会に、生成AIについての感想を複数の方から聞いた。

Aさん「ChatGPTがよく間違えるので、最近はGeminiを使っているんです」
Bさん「Claudeはあまり当てにならないので、ChatGPTばかり使っています」
Cさん「ハルシネーションの心配はありますが、数字は間違いないですよ」

このうちCさんには、僕自身の経験から口を挟んだ。

「数字も間違えることがあるよ。間違えるというより、桁を取り違えたり、表の数字を読み飛ばしたりすることもある。盲信は禁物だと思う。」

けれどもAさんとBさんには、僕自身は違う印象を持っていたものの、あえて反論しなかった。

そもそも、お二人が「いつ」使ったのかが分からない。これが一番大きいと思う。生成AIの進歩はかなり速いからだ。一口にChatGPTと言っても、どのモデルを使ったのかも分からない。何をさせようとして、どのようなプロンプトを書いたときに、そういう結果になったのか。さらに言えば、間違いに気づかないまま正しいと思い込んだケースがどれだけあったのか。これは尋ねようがないし、おそらく本人にも分からない。

三人がそれぞれそういう場面に出くわしたこと自体は、事実だろう。しかし、そこからChatGPT、Gemini、Claudeのどれが正確で、どれが当てにならないかを結論づけることは難しい。

普通の商品やサービスなら、品質や特徴は比較的安定しているので、複数の利用者が同じような長所や欠点について語り、その評価をある程度共有することができる。ところが生成AIは、同じ名前のサービスでも、使う時期やモデルによって中身が違う。何をさせるか、どう指示するかによっても結果は大きく変わる。しかも、利用者が誤りを見抜けるかどうかによって、そのAIに対する評価まで変わってしまう。

生成AIを評価することの難しさは、まさにそこにあるのだ、と飲みながら(はい、いずれも飲み会でした)つくづく思った。