機械学習モデルの結果を英語で説明する表現|精度・過学習・解釈性

機械学習モデルの結果を英語で説明する表現|精度・過学習・解釈性

機械学習モデルの結果を英語で説明するときは、accuracyやR-squaredなどの指標を読み上げるだけでは十分ではありません。「どのデータで評価したか」「学習データとの差はどうか」「どこまで一般化できるか」「予測根拠をどのように確認したか」まで説明すると、モデルの実用性が伝わります。

本記事では、回帰・分類モデルの精度、交差検証、過学習、データリーク、解釈性、不確かさ、適用領域を英語で説明する表現を、日本語訳とともに紹介します。国際学会、海外との共同研究、技術報告、論文発表でそのまま応用できる構成です。

この記事で扱う内容

  • モデル結果を説明する基本構成
  • MAE、RMSE、R²、accuracy、precision、recall、F1の表現
  • 訓練・検証・テストデータの説明
  • 過学習、未学習、データリークの伝え方
  • 特徴量重要度、SHAP、部分依存性と解釈上の注意
  • 不確かさ、外挿、適用領域の説明
  • 材料科学の発表例と質疑応答

モデル結果を英語で説明する基本の順番

機械学習の結果は、次の5点を順番に説明すると伝わりやすくなります。

  1. Task:何を予測・分類したか
  2. Data:どのデータで学習・評価したか
  3. Metric:どの評価指標を使ったか
  4. Result:どの程度の性能だったか
  5. Interpretation:結果をどう解釈し、どこまで利用できるか

We developed a regression model to predict tensile strength from composition and processing conditions. The model was evaluated on a held-out test set using mean absolute error and R-squared. It achieved an MAE of 4.8 MPa and an R² of 0.87. The results indicate that the model captures the main trend within the represented processing range.

組成とプロセス条件から引張強度を予測する回帰モデルを構築しました。独立したテストデータに対して平均絶対誤差と決定係数で評価したところ、MAEは4.8 MPa、R²は0.87でした。この結果は、データがカバーするプロセス範囲内でモデルが主要な傾向を捉えていることを示しています。

The model is highly accurate.と結論だけを述べるより、評価データ、指標、数値、利用範囲を一緒に示す方が、再現性と説得力のある説明になります。モデル評価では、目的に合う指標と検証方法を選ぶことが重要です。[1、2]

精度を説明する基本表現

性能が良かったと伝える

  • The model showed strong predictive performance on the test set.
    モデルはテストデータで高い予測性能を示しました。
  • The predictions were in good agreement with the measured values.
    予測値は測定値とよく一致しました。
  • The model reproduced the overall trend across the evaluated range.
    モデルは評価範囲全体の傾向を再現しました。
  • The prediction error was sufficiently small for candidate screening.
    予測誤差は候補スクリーニングの用途には十分小さいものでした。
  • The model met the predefined performance criterion.
    モデルは事前に定めた性能基準を満たしました。

結果を控えめに評価する

  • The model showed moderate predictive performance.
    モデルは中程度の予測性能を示しました。
  • The model captured the general trend, although some individual predictions showed large errors.
    一部の個別予測には大きな誤差がありましたが、モデルは全体的な傾向を捉えました。
  • The performance is promising for prioritization, but further validation is needed before quantitative use.
    優先順位づけへの利用には有望ですが、定量的な利用には追加検証が必要です。
  • The current model should be regarded as a screening tool rather than a substitute for measurement.
    現在のモデルは、測定の代替ではなくスクリーニングツールとして位置づけるべきです。

モデル間の差を説明する

  • Model A outperformed Model B in terms of MAE.
    MAEの点では、モデルAがモデルBを上回りました。
  • The random forest reduced the RMSE by 18% compared with the linear baseline.
    ランダムフォレストは線形ベースラインと比べてRMSEを18%低減しました。
  • The two models achieved comparable performance within the variability of cross-validation.
    交差検証のばらつきを考慮すると、2つのモデルの性能は同程度でした。
  • The improvement was consistent across all five folds.
    性能向上は5つの分割すべてで一貫していました。
  • The more complex model provided only a marginal improvement.
    より複雑なモデルによる改善はわずかでした。

回帰モデルの評価指標を説明する

MAE:Mean Absolute Error

The mean absolute error was 3.2 degrees Celsius, meaning that the predictions differed from the measured values by 3.2 degrees on average.

平均絶対誤差は3.2℃でした。これは、予測値と測定値の差が平均で3.2℃だったことを意味します。

MAEは元の目的変数と同じ単位なので、実務上の意味を説明しやすい指標です。

  • A lower MAE indicates smaller errors on average.
    MAEが小さいほど、平均的な誤差が小さいことを示します。
  • The MAE was below our allowable experimental variation.
    MAEは許容する実験変動より小さい値でした。
  • The error corresponds to approximately 6% of the target range.
    この誤差は目的値の範囲のおよそ6%に相当します。

RMSE:Root Mean Squared Error

The root mean squared error was 5.6 MPa. Because RMSE gives greater weight to large errors, we examined the corresponding outliers separately.

二乗平均平方根誤差は5.6 MPaでした。RMSEは大きな誤差に強く影響されるため、対応する外れ値を別途確認しました。

  • The RMSE was larger than the MAE, suggesting that a small number of samples had relatively large errors.
    RMSEはMAEより大きく、少数の試料で比較的大きな誤差があったことが示唆されます。
  • The normalized RMSE was 0.09.
    正規化RMSEは0.09でした。

R²:Coefficient of Determination

The model achieved an R-squared value of 0.82 on the test set, indicating that it explained a substantial proportion of the observed variation.

モデルはテストデータでR²=0.82を示し、観測された変動の大部分を説明しました。

R² of 0.82 means that the model is 82% accurate.という説明は避けた方がよいでしょう。R²は「正解率」ではなく、目的変数の変動をモデルがどの程度説明したかを表す指標です。

  • A high R² does not necessarily imply small absolute errors.
    R²が高くても、絶対的な誤差が小さいとは限りません。
  • We therefore report both R² and MAE.
    そのため、R²とMAEの両方を報告します。
  • The test-set R² was lower than the cross-validation estimate.
    テストデータのR²は交差検証による推定値より低い値でした。

分類モデルの評価指標を説明する

Accuracy:正解率

The classifier achieved an accuracy of 91% on the test set.

分類モデルはテストデータで91%の正解率を達成しました。

クラス比率が偏っている場合、accuracyだけでは性能を適切に示せないことがあります。

Because the classes were imbalanced, accuracy alone was not sufficient. We also evaluated precision, recall, and the F1 score.

クラスに偏りがあったため、accuracyだけでは十分ではありませんでした。そこでprecision、recall、F1スコアも評価しました。

Precision:適合率

  • The precision was 0.88, meaning that 88% of the samples predicted as positive were actually positive.
    適合率は0.88でした。陽性と予測した試料の88%が実際に陽性だったことを意味します。
  • High precision is important when false positives lead to costly experiments.
    偽陽性によって高コストの実験が発生する場合、高い適合率が重要です。

Recall:再現率

  • The recall was 0.93, indicating that the model identified 93% of the actual positive samples.
    再現率は0.93で、実際の陽性試料の93%をモデルが検出したことを示します。
  • We prioritized recall because missing a promising candidate was considered more costly than testing an additional sample.
    有望候補を見逃すことは追加試料を評価することより損失が大きいと考え、再現率を重視しました。

F1 scoreとROC-AUC

  • The F1 score was 0.90, reflecting a good balance between precision and recall.
    F1スコアは0.90で、適合率と再現率のバランスが良好でした。
  • The model achieved an ROC-AUC of 0.94.
    モデルのROC-AUCは0.94でした。
  • We selected the decision threshold based on the required trade-off between false positives and false negatives.
    偽陽性と偽陰性の必要なバランスに基づいて判定閾値を選びました。

訓練・検証・テスト結果を英語で説明する

English 日本語
We used 70% of the data for training, 15% for validation, and 15% for testing. データの70%を学習、15%を検証、15%をテストに使用しました。
The validation set was used for hyperparameter selection. 検証データはハイパーパラメータ選択に使用しました。
The test set was held out until the final evaluation. テストデータは最終評価まで使用せずに分離しました。
No test-set information was used during model development. モデル開発中にテストデータの情報は使用していません。
We used five-fold cross-validation to estimate generalization performance. 汎化性能を推定するため、5分割交差検証を使用しました。
The reported score is the mean across five folds, and the error bar represents one standard deviation. 報告値は5分割の平均で、エラーバーは1標準偏差を表します。

交差検証の結果を伝える

The cross-validated MAE was 4.1 plus or minus 0.6 MPa. The relatively small variation across folds suggests that the model performance was stable with respect to the data split.

交差検証によるMAEは4.1±0.6 MPaでした。分割間のばらつきが比較的小さいことから、データ分割に対してモデル性能が安定していたことが示されます。

  • The performance varied considerably across folds.
    分割によって性能が大きく変動しました。
  • This variability appears to reflect differences among material families.
    このばらつきは材料ファミリー間の違いを反映していると考えられます。
  • We used grouped cross-validation to prevent closely related samples from appearing in both training and validation sets.
    類似試料が学習データと検証データの両方に入ることを防ぐため、グループ化交差検証を用いました。

ハイパーパラメータを調整しながら同じテストデータを繰り返し参照すると、その情報がモデル選択に入り込み、汎化性能を楽観的に評価する可能性があります。scikit-learnの公式文書でも、テストデータに基づく調整による情報漏れが説明されています。[1]

過学習を英語で説明する

overfittingは、学習データにはよく適合する一方、未知データへの性能が低下している状態です。

The training error continued to decrease, whereas the validation error began to increase after 100 iterations. This divergence indicates overfitting.

学習誤差は低下し続けましたが、検証誤差は100回の反復後から増加しました。この乖離は過学習を示しています。

過学習の兆候を説明する

  • The model performed substantially better on the training set than on the validation set.
    モデルは検証データより学習データで大幅に高い性能を示しました。
  • A large train–validation gap was observed.
    学習性能と検証性能の間に大きな差が見られました。
  • The model appears to have learned noise or dataset-specific patterns.
    モデルはノイズまたはそのデータセット固有のパターンを学習した可能性があります。
  • Increasing model complexity improved the training score but reduced the validation score.
    モデルを複雑にすると学習スコアは向上しましたが、検証スコアは低下しました。

対策を説明する

  • We reduced overfitting through regularization and early stopping.
    正則化と早期終了によって過学習を抑えました。
  • We simplified the model and removed unstable features.
    モデルを単純化し、不安定な特徴量を除きました。
  • Data augmentation improved validation performance.
    データ拡張によって検証性能が向上しました。
  • We selected the model at the minimum validation error rather than the minimum training error.
    学習誤差が最小のモデルではなく、検証誤差が最小のモデルを選びました。
  • Nested cross-validation was used to separate hyperparameter tuning from performance estimation.
    ハイパーパラメータ調整と性能推定を分離するため、ネストした交差検証を用いました。

学習曲線や検証曲線では、学習スコアと検証スコアの変化を比較することで、過学習と未学習を判断できます。[3]

未学習を英語で説明する

underfittingは、モデルがデータ内の主要な関係を十分に表現できず、学習データでも検証データでも性能が低い状態です。

Both the training and validation errors remained high, suggesting that the model was underfitting the data.

学習誤差と検証誤差の両方が高いままであり、モデルがデータに十分適合していないことが示唆されました。

  • The linear model was too simple to capture the nonlinear relationship.
    線形モデルは非線形な関係を捉えるには単純すぎました。
  • Adding physically meaningful interaction terms improved both training and validation performance.
    物理的に意味のある交互作用項を加えることで、学習・検証性能の両方が向上しました。
  • Excessive regularization reduced the model’s ability to learn the signal.
    過度の正則化によって、モデルが信号を学習する能力が低下しました。

データリークを英語で説明する

data leakageは、予測時には利用できない情報や、評価データ由来の情報が学習工程に入り込むことです。

  • We identified data leakage caused by preprocessing the full dataset before splitting.
    データ分割前に全データを前処理したことで生じたデータリークを確認しました。
  • All preprocessing steps were fitted using the training data only.
    すべての前処理は学習データのみを使って適合させました。
  • Replicate measurements from the same specimen were kept in the same fold.
    同一試料の反復測定値は同じ分割にまとめました。
  • The original random split produced an overly optimistic score.
    当初のランダム分割では、過度に楽観的なスコアが得られていました。
  • After removing the leakage, the test accuracy decreased from 96% to 84%, providing a more realistic estimate.
    リークを除くとテスト正解率は96%から84%に低下し、より現実的な性能推定となりました。

解釈性を英語で説明する

モデルの解釈性は、単に「特徴量重要度を計算した」と述べるだけでなく、何を説明しているかを明確にすることが重要です。

モデル全体を説明する

  • We used feature importance to examine the model’s global behavior.
    モデル全体の挙動を調べるため、特徴量重要度を使用しました。
  • Processing temperature was the most influential feature in the model.
    モデルではプロセス温度が最も影響の大きい特徴量でした。
  • The partial dependence plot showed a nonlinear relationship between curing temperature and predicted strength.
    部分依存プロットは、硬化温度と予測強度の間に非線形な関係があることを示しました。
  • The model relied primarily on composition-related descriptors.
    モデルは主に組成関連の記述子に依存していました。

個別の予測を説明する

  • We used SHAP values to explain individual predictions.
    個々の予測を説明するため、SHAP値を使用しました。
  • For this sample, the high predicted value was mainly associated with features A and B.
    この試料の高い予測値は、主に特徴量AとBに関連していました。
  • The local explanation shows how each feature shifted the prediction relative to the baseline.
    局所的な説明は、各特徴量が基準値から予測をどのように変化させたかを示します。

因果関係と区別する

Feature importance reflects how the model uses the variables for prediction. It should not be interpreted as direct evidence of a causal relationship.

特徴量重要度は、モデルが予測に変数をどのように利用しているかを表します。因果関係の直接的な証拠として解釈すべきではありません。

  • The observed association is consistent with the proposed mechanism, but does not establish causality.
    観測された関連は提案メカニズムと整合しますが、因果関係を確立するものではありません。
  • Correlated features may share or redistribute importance.
    相関する特徴量の間では重要度が分散または再配分されることがあります。
  • We confirmed the interpretation against domain knowledge and additional experiments.
    専門知識と追加実験によって解釈の妥当性を確認しました。

材料科学の説明可能AIでは、グローバルな説明と個別予測の説明、モデル固有の解釈性と事後的な説明を区別する考え方が整理されています。また、説明自体が誤解を招く可能性があるため、説明方法の評価も必要です。[4]

不確かさと適用領域を説明する

予測の不確かさ

  • The shaded region represents the 95% prediction interval.
    網掛け部分は95%予測区間を表します。
  • Prediction uncertainty was low near the training samples and increased in sparsely sampled regions.
    予測の不確かさは学習試料の近傍では小さく、データの少ない領域で大きくなりました。
  • We used ensemble variance as an approximate measure of model uncertainty.
    モデル不確かさの近似指標として、アンサンブル間の分散を用いました。
  • The point estimate should be interpreted together with its uncertainty.
    点推定値は不確かさとあわせて解釈する必要があります。

内挿と外挿

  • Most test samples were within the interpolation range of the training data.
    テスト試料の大部分は学習データの内挿範囲内にありました。
  • The model was less reliable for extrapolative predictions.
    外挿的な予測では、モデルの信頼性が低下しました。
  • This candidate lies outside the distribution represented in the training set.
    この候補は、学習データで表現された分布の外側にあります。
  • Additional measurements are required before applying the model to this composition family.
    この組成ファミリーにモデルを適用する前に、追加測定が必要です。

適用領域

We defined the applicability domain using compositional similarity and the range of processing conditions. Predictions outside this domain were flagged for review rather than used automatically.

組成の類似性とプロセス条件の範囲を用いて適用領域を定義しました。この領域外の予測は自動的に利用せず、確認対象としてフラグを付けました。

材料科学では、モデル全体の平均性能だけでなく、どの材料領域で予測が信頼できるかを特定することが重要です。適用領域の評価により、局所的に良好な領域と信頼性の低い領域を分けて考えられます。[5]

グラフを使って結果を説明する表現

Parity plot:予測値と実測値の比較

  • This parity plot compares the predicted and measured values.
    このパリティプロットは予測値と測定値を比較しています。
  • Points close to the diagonal indicate accurate predictions.
    対角線に近い点は、予測が正確であることを示します。
  • The model tended to underestimate the highest values.
    モデルは高い値を過小評価する傾向がありました。
  • No systematic bias was observed across the central range.
    中央の範囲では系統的な偏りは見られませんでした。

Residual plot:残差プロット

  • The residuals were centered around zero.
    残差はゼロ付近を中心に分布していました。
  • The residual variance increased with the target value.
    目的値が大きくなるにつれて残差の分散が増加しました。
  • The residual pattern suggests that the model does not fully capture the nonlinear trend.
    残差のパターンから、モデルが非線形な傾向を十分に捉えていないことが示唆されます。

Confusion matrix:混同行列

  • The confusion matrix summarizes correct and incorrect classifications for each class.
    混同行列は、各クラスの正分類と誤分類をまとめたものです。
  • Most errors occurred between classes B and C.
    誤分類の大部分はクラスBとCの間で生じました。
  • The model rarely misclassified negative samples as positive.
    モデルが陰性試料を陽性と誤分類することはほとんどありませんでした。

グラフの増減、比較、誤差を説明する一般表現は、英語でグラフを説明する表現集でも詳しく紹介しています。

材料科学での説明例

回帰モデルの発表例

We trained a gradient-boosting model to predict the ionic conductivity of solid electrolytes from compositional and structural descriptors. Hyperparameters were selected using nested cross-validation, and final performance was evaluated on a held-out material-family test set. The model achieved an MAE of 0.31 log units and an R² of 0.79. The test error was slightly higher than the cross-validation estimate, but the difference was within one standard deviation. SHAP analysis indicated that volume-related and electronegativity-related descriptors had the largest influence on the predictions. These associations were consistent with known transport considerations, although we do not interpret them as direct causal evidence. Prediction uncertainty increased for compositions distant from the training data, so those candidates were assigned to experimental validation rather than automatic ranking.

日本語訳:組成・構造記述子から固体電解質のイオン伝導度を予測する勾配ブースティングモデルを学習しました。ネストした交差検証でハイパーパラメータを選択し、最終性能は材料ファミリー単位で分離したテストデータで評価しました。モデルのMAEは0.31 log単位、R²は0.79でした。テスト誤差は交差検証の推定値よりわずかに大きいものの、その差は1標準偏差の範囲内でした。SHAP解析では、体積と電気陰性度に関する記述子が予測に大きく影響していました。これらの関連は既知の輸送に関する知見と整合しますが、直接的な因果関係とは解釈していません。学習データから離れた組成では予測不確かさが増えたため、それらの候補は自動順位づけではなく実験検証の対象としました。

分類モデルの発表例

We developed a classifier to identify compositions likely to form a single phase. Because only 18% of the samples belonged to the positive class, we evaluated precision, recall, and the F1 score in addition to accuracy. The model achieved a precision of 0.81 and a recall of 0.89 on the test set. We selected the decision threshold to prioritize recall, because missing a promising composition was more costly than conducting a limited number of additional experiments.

日本語訳:単相を形成する可能性が高い組成を識別する分類モデルを構築しました。陽性クラスは試料の18%のみだったため、正解率に加えて適合率、再現率、F1スコアを評価しました。テストデータで適合率0.81、再現率0.89を達成しました。有望な組成を見逃す損失が、少数の追加実験を行う損失より大きいため、再現率を重視して判定閾値を選びました。

質疑応答で使える英語表現

「なぜこの評価指標を選んだのですか」

We selected MAE because it has the same unit as the target property and is straightforward to interpret in relation to experimental variation. We also report R² to describe how well the model captures variation across the dataset.

MAEは目的物性と同じ単位であり、実験変動との関係を解釈しやすいため選びました。データセット全体の変動をモデルがどの程度捉えているかを示すため、R²も報告しています。

「モデルは過学習していませんか」

We checked for overfitting by comparing training and validation learning curves. The final model showed a small and stable performance gap. We also confirmed its performance on a test set that was not used for model selection.

学習曲線と検証曲線を比較して過学習を確認しました。最終モデルでは性能差が小さく安定していました。また、モデル選択に使用していないテストデータでも性能を確認しました。

「ランダム分割でよいのですか」

A random split would place closely related compositions in both sets and could overestimate generalization performance. We therefore grouped the data by composition family before splitting.

ランダム分割では類似組成が両方のデータに入り、汎化性能を過大評価する可能性があります。そのため、分割前に組成ファミリーでデータをグループ化しました。

「特徴量重要度は因果関係を表しますか」

No. Feature importance shows how the trained model uses the available variables. It can support hypothesis generation, but causal interpretation requires an appropriate experimental design or causal analysis.

いいえ。特徴量重要度は、学習済みモデルが利用可能な変数をどのように使うかを示します。仮説生成には役立ちますが、因果的な解釈には適切な実験計画または因果解析が必要です。

「このモデルを新しい材料系に使えますか」

The current evidence supports interpolation within the represented material space. Application to a new material family should be treated as extrapolation and requires additional validation.

現在の検証結果が支持しているのは、データで表現された材料空間内の内挿です。新しい材料ファミリーへの適用は外挿として扱い、追加検証を行う必要があります。

断定の強さを調整する英語表現

強さ 英語表現 使い方
強い The results demonstrate that… 複数の検証で明確に支持された場合
やや強い The results indicate that… 結果が一定の方向を示している場合
中程度 The results suggest that… 合理的な示唆はあるが追加検証が必要な場合
限定的 The results are consistent with… 既存仮説と矛盾しないことを述べる場合
仮説 One possible explanation is that… 考えられる説明を提示する場合

結果に合った動詞を使うことで、過度な断定を避けながら、必要以上に否定的にならずに説明できます。

そのまま使える結果説明テンプレート

We developed a [regression/classification] model to ______.

The dataset consisted of ______ samples, which were divided into ______.

We used ______ as the primary evaluation metric because ______.

The model achieved ______ on the held-out test set.

Compared with the baseline, the model ______.

The difference between training and validation performance was ______, indicating ______.

Interpretability analysis showed that ______.

These results should be interpreted within ______.

The next step is to validate ______.

短い完成例

We developed a regression model to predict thermal conductivity. The model achieved an MAE of 0.18 W m⁻¹ K⁻¹ on the held-out test set, representing a 22% improvement over the linear baseline. Training and validation errors were similar, and no substantial overfitting was observed. Feature analysis suggested that density and porosity were the most influential predictors. The model is intended for candidate screening within the processing range represented by the training data.

よくある英語表現の修正

避けたい表現 より自然・正確な表現
The accuracy was good. The model achieved an MAE of 4.2 MPa on the test set.
R² was 80% accuracy. The test-set R² was 0.80.
The model completely avoided overfitting. We found no substantial evidence of overfitting under the current validation scheme.
Feature A caused the prediction. Feature A made a large contribution to the model prediction.
The model can predict unknown materials. The model showed predictive performance within the evaluated material space.
The AI discovered the mechanism. The analysis generated a hypothesis that was consistent with the experimental trend.

関連する英語記事

まとめ

機械学習モデルの結果を英語で説明するときは、単一の精度指標ではなく、評価データ、指標を選んだ理由、訓練・検証性能の差、解釈性、不確かさ、適用領域を一続きで示すことが大切です。

基本形として、次の表現を覚えておくと便利です。

The model achieved [metric] on the held-out test set. Training and validation performance were comparable, suggesting limited overfitting. The interpretation applies within the domain represented by the training data.

モデルは独立したテストデータで[指標]を達成しました。学習性能と検証性能は同程度であり、過学習は限定的であることが示唆されます。この解釈は、学習データで表現された領域内に適用されます。

この基本形に、研究目的、ベースラインとの比較、特徴量の解釈、次の実験検証を加えると、国際学会や海外会議でも説得力のある説明になります。

参考文献

  1. scikit-learn developers, “Cross-validation: evaluating estimator performance.”
  2. scikit-learn developers, “Metrics and scoring: quantifying the quality of predictions.”
  3. scikit-learn developers, “Validation curves: plotting scores to evaluate models.”
  4. Zhong, X. et al., “Explainable machine learning in materials science,” npj Computational Materials, 8, 204 (2022). DOI: 10.1038/s41524-022-00884-7.
  5. Sutton, C. et al., “Identifying domains of applicability of machine learning models for materials science,” Nature Communications, 11, 4428 (2020). DOI: 10.1038/s41467-020-17112-9.
  6. National Institute of Standards and Technology, “AI Measurement and Evaluation.”

コメントする

メールアドレスが公開されることはありません。 ※ が付いている欄は必須項目です

CAPTCHA


上部へスクロール