統計解析結果を英語で書く方法|p値・信頼区間・効果量の例文

統計解析結果を英語で書く方法|p値・信頼区間・効果量の例文

統計解析結果を英語で書くとき、p < 0.05だけを示しても、差の大きさや推定の精度は伝わりません。読み手が知りたいのは、「どの程度の差があり」「その推定にはどのくらいの不確かさがあり」「研究上・実務上どのような意味を持つか」です。

本記事では、p値、信頼区間、効果量を英語で報告する基本形から、t検定、ANOVA、相関、回帰、比率比較、多重比較の例文までを日本語訳付きで紹介します。論文、国際学会、研究報告書、海外との技術会議に応用できる表現をまとめました。

この記事でできるようになること

  • 推定値・信頼区間・p値を一続きで報告する
  • 統計的有意性と実質的な重要性を区別する
  • 「有意差がなかった」を正確に表現する
  • 効果量を研究上の意味と結びつける
  • t検定、ANOVA、相関、回帰の結果を書く
  • 材料研究の統計結果を英語で説明する

統計解析結果を書く基本の順番

統計結果は、次の順で書くと読み手が解釈しやすくなります。

  1. Descriptive result:各群の値や差
  2. Effect estimate:平均差、比、回帰係数など
  3. Confidence interval:推定の範囲
  4. p-value:仮説検定の結果
  5. Interpretation:研究上・実務上の意味

The modified formulation increased the mean adhesion strength from 18.4 to 22.1 MPa. The mean difference was 3.7 MPa (95% CI, 1.8 to 5.6 MPa; p = 0.001), corresponding to a large standardized effect (Cohen’s d = 1.02). The improvement also exceeded the predefined practical threshold of 2 MPa.

改良配合によって平均接着強度は18.4 MPaから22.1 MPaへ増加しました。平均差は3.7 MPa(95%信頼区間:1.8~5.6 MPa、p=0.001)で、標準化効果量は大きい値でした(Cohenのd=1.02)。この改善は、事前に定めた実用上の閾値である2 MPaも上回りました。

この書き方なら、有意かどうかだけでなく、差の方向、大きさ、推定精度、実務上の意味が分かります。統計報告のガイドラインでも、記述統計、効果の推定値、信頼区間を示し、p値だけに依存しない報告が推奨されています。[1、2]

p値を英語で報告する基本表現

統計的に有意だった場合

  • The difference was statistically significant (p = 0.012).
    差は統計的に有意でした(p=0.012)。
  • Group A had a significantly higher mean value than Group B (p = 0.004).
    グループAの平均値はグループBより有意に高い値でした(p=0.004)。
  • A statistically significant association was observed between temperature and yield (p = 0.021).
    温度と収率の間に統計的に有意な関連が認められました(p=0.021)。
  • The null hypothesis was rejected at the prespecified 5% significance level.
    事前に定めた5%の有意水準で帰無仮説を棄却しました。

統計的に有意でなかった場合

  • The difference was not statistically significant (p = 0.18).
    差は統計的に有意ではありませんでした(p=0.18)。
  • The result did not reach the prespecified level of statistical significance.
    結果は事前に定めた統計的有意水準に達しませんでした。
  • We did not reject the null hypothesis.
    帰無仮説は棄却されませんでした。
  • The data did not provide sufficient evidence of a difference between the groups.
    群間差を示す十分な証拠は、今回のデータから得られませんでした。

There was no difference.やThe groups were the same.は、非有意という結果から同等性まで結論づける表現です。差が検出されなかった場合は、did not provide sufficient evidenceやwas not statistically significantと書きます。同等性を主張するには、同等性試験や非劣性試験など、目的に合った設計が必要です。

p値を正確に示す

  • p = 0.032:可能であれば実際のp値を示す
  • p < 0.001:非常に小さい場合に不等号で示す
  • p = 0.000:p値は厳密にゼロではないため避ける
  • p = NS:情報量が少ないため、具体的なp値を示す方がよい

The exact p-value was 0.047.

正確なp値は0.047でした。

報告規程に別の指定がなければ、p < 0.05と閾値だけを書くより、p = 0.047のように実際の値を示す方が情報を保てます。SAMPLガイドラインも、可能な場合はp値を等号で報告することを推奨しています。[2、3]

p値が意味すること・意味しないこと

p値は、指定した統計モデルと帰無仮説の下で、観測された統計量と同じか、それ以上に極端な結果が得られる確率を表します。

適切な英語表現

Under the null hypothesis and the specified statistical model, a result at least as extreme as the observed one would occur with probability p = 0.03.

帰無仮説と指定した統計モデルの下では、観測結果と同程度以上に極端な結果が得られる確率はp=0.03です。

避けたい解釈

誤解を招く表現 問題点
There is only a 3% probability that the null hypothesis is true. p値は帰無仮説が真である確率ではありません。
The result has a 97% probability of being correct. p値から結論が正しい確率は計算できません。
The result occurred by chance. p値は「偶然だけで生じた確率」そのものではありません。
A smaller p-value means a larger effect. p値は効果の大きさを直接表しません。
p > 0.05 proves that there is no effect. 非有意は効果がないことの証明ではありません。

米国統計学会の声明でも、p値は仮説が真である確率や、効果の大きさ・結果の重要性を測るものではないと整理されています。[1]

信頼区間を英語で報告する

基本表現

  • The mean difference was 2.4 MPa (95% CI, 0.8 to 4.0 MPa).
    平均差は2.4 MPaでした(95%信頼区間:0.8~4.0 MPa)。
  • The 95% confidence interval ranged from 0.8 to 4.0 MPa.
    95%信頼区間は0.8~4.0 MPaでした。
  • The estimate was 1.35, with a 95% CI of 1.10 to 1.66.
    推定値は1.35で、95%信頼区間は1.10~1.66でした。
  • The confidence interval was relatively narrow, indicating a precise estimate.
    信頼区間は比較的狭く、精度の高い推定であることを示しています。
  • The wide confidence interval reflects substantial uncertainty.
    広い信頼区間は、大きな不確かさを反映しています。

信頼区間から結果の範囲を説明する

The estimated increase was 3.0 MPa, and the 95% confidence interval ranged from 0.2 to 5.8 MPa. The data are therefore compatible with effects ranging from a small to a substantial improvement.

推定された増加量は3.0 MPaで、95%信頼区間は0.2~5.8 MPaでした。したがって今回のデータは、小さな改善から大きな改善までの効果と整合します。

信頼区間は、単にゼロを含むかどうかだけでなく、研究上重要な値を含むかを確認します。

  • The interval excluded zero but included effects below the practical threshold.
    区間はゼロを含みませんでしたが、実用上の閾値を下回る効果を含んでいました。
  • The entire interval was above the predefined minimum relevant difference.
    区間全体が、事前に定めた最小の実質的差を上回っていました。
  • The interval included both negligible and practically important effects.
    区間には、無視できる効果と実用上重要な効果の両方が含まれていました。
  • More data are needed to estimate the effect with useful precision.
    実用的な精度で効果を推定するには、追加データが必要です。

95%信頼区間の解釈

頻度論的な95%信頼区間は、「今回得られた区間に真の値が95%の確率で入る」という事後確率を直接表すものではありません。同じ手順で標本抽出と区間推定を繰り返したとき、作成される区間の約95%が真の母数を含むように設計された手法です。

英語では、無理に確率へ置き換えず、次のように書くと安全です。

The estimated effect was 2.4 MPa, with a 95% confidence interval from 0.8 to 4.0 MPa.

効果量を英語で報告する

効果量は、差や関連の大きさを示します。目的に応じて、元の単位で表す効果量と、標準化効果量を使い分けます。

平均差:Mean difference

  • The mean difference between the groups was 4.2 MPa.
    群間の平均差は4.2 MPaでした。
  • The treatment increased the mean response by 12 percentage points.
    処理によって平均応答が12ポイント増加しました。
  • The adjusted mean difference was 3.6 MPa after controlling for batch effects.
    バッチ効果を調整した平均差は3.6 MPaでした。

元の単位による平均差は、実験誤差や仕様値と直接比較できるため、材料研究・製造研究では特に有用です。

標準化平均差:Cohen’s d

  • The standardized mean difference was Cohen’s d = 0.72.
    標準化平均差はCohenのd=0.72でした。
  • The effect was moderate in standardized terms (d = 0.58).
    標準化した基準では中程度の効果でした(d=0.58)。
  • The confidence interval for Cohen’s d was wide because of the small sample size.
    サンプル数が少ないため、Cohenのdの信頼区間は広いものでした。

「0.2=小、0.5=中、0.8=大」という目安は分野横断的な便宜的基準です。研究上の重要性は、分野固有の変動、測定精度、要求仕様とあわせて判断します。

相関係数:Correlation coefficient

  • Temperature was positively correlated with conversion (Pearson’s r = 0.64, 95% CI, 0.38 to 0.81; p < 0.001).
    温度は転化率と正の相関を示しました(Pearsonのr=0.64、95%信頼区間:0.38~0.81、p<0.001)。
  • The correlation was weak and uncertain (r = 0.18, 95% CI, -0.12 to 0.45).
    相関は弱く、不確かさが大きい結果でした(r=0.18、95%信頼区間:−0.12~0.45)。
  • The association should not be interpreted as causal.
    この関連を因果関係として解釈すべきではありません。

オッズ比・リスク比

  • The odds of failure were lower in Group A than in Group B (odds ratio = 0.42, 95% CI, 0.21 to 0.83; p = 0.013).
    故障のオッズはグループAでグループBより低い値でした(オッズ比=0.42、95%信頼区間:0.21~0.83、p=0.013)。
  • The relative risk was 1.28 (95% CI, 1.05 to 1.56).
    相対リスクは1.28でした(95%信頼区間:1.05~1.56)。
  • An odds ratio should not be described as a risk ratio when the event is common.
    事象が一般的な場合、オッズ比をリスク比として説明すべきではありません。

ANOVAの効果量

  • The factor explained 18% of the variance (partial eta squared = 0.18).
    この因子は分散の18%を説明しました(偏η²=0.18)。
  • The effect size was reported as omega squared to reduce small-sample bias.
    小標本でのバイアスを軽減するため、効果量をω²で報告しました。

統計的有意性と実質的な重要性を区別する

statistically significantは「統計的に有意」、practically meaningfulは「実用上意味がある」、scientifically importantは「科学的に重要」を表します。

有意だが、効果が小さい

The difference was statistically significant (p < 0.001), but the mean improvement was only 0.3%, which was below the predefined practical threshold of 2%.

差は統計的に有意でした(p<0.001)が、平均改善量は0.3%にすぎず、事前に定めた実用上の閾値2%を下回りました。

非有意だが、重要な効果を除外できない

The result was not statistically significant (p = 0.11); however, the 95% confidence interval included improvements that would be practically important. The study may not have estimated the effect with sufficient precision.

結果は統計的に有意ではありませんでした(p=0.11)。ただし、95%信頼区間には実用上重要な改善量も含まれていました。この研究では、効果を十分な精度で推定できていない可能性があります。

差が小さいことを比較的精密に示した

The estimated difference was close to zero, and the confidence interval excluded the predefined meaningful effect. These results suggest that any remaining difference is likely to be small within the tested conditions.

推定差はゼロに近く、信頼区間は事前に定めた意味のある効果を除外していました。この結果は、試験条件内で残る差が小さい可能性を示しています。

統計的有意かどうかの二分法だけで結論を決めず、効果の大きさ、不確かさ、研究設計、先行知見をあわせて解釈する考え方は、米国統計学会の声明とその後の提言でも重視されています。[1、4]

t検定の結果を英語で書く

独立2群のt検定

The mean tensile strength was higher for Formulation A than for Formulation B (52.4 ± 4.1 vs. 47.8 ± 3.7 MPa). The mean difference was 4.6 MPa (95% CI, 1.5 to 7.7 MPa; t(22) = 3.08, p = 0.006; Cohen’s d = 1.18).

平均引張強度は配合Bより配合Aで高い値でした(52.4±4.1 MPa対47.8±3.7 MPa)。平均差は4.6 MPaでした(95%信頼区間:1.5~7.7 MPa、t(22)=3.08、p=0.006、Cohenのd=1.18)。

Welchのt検定

Because the group variances differed, we used Welch’s t-test. The estimated mean difference was 3.1 MPa (95% CI, 0.4 to 5.8 MPa; t(15.7) = 2.44, p = 0.027).

群間で分散が異なったため、Welchのt検定を使用しました。推定平均差は3.1 MPaでした(95%信頼区間:0.4~5.8 MPa、t(15.7)=2.44、p=0.027)。

対応のあるt検定

After treatment, the mean response increased by 2.8 units (95% CI, 1.1 to 4.5 units; paired t(11) = 3.62, p = 0.004).

処理後、平均応答は2.8単位増加しました(95%信頼区間:1.1~4.5単位、対応のあるt検定:t(11)=3.62、p=0.004)。

ANOVAの結果を英語で書く

一元配置分散分析

Mean conductivity differed among the four formulations (one-way ANOVA, F(3, 36) = 6.42, p = 0.001; partial η² = 0.35).

平均伝導度は4つの配合間で異なりました(一元配置分散分析:F(3, 36)=6.42、p=0.001、偏η²=0.35)。

ANOVAの有意結果は、「少なくともどこかの群間に差がある」ことを示します。どの群が異なるかは、事後比較とともに報告します。

Tukey-adjusted pairwise comparisons showed that Formulation D had higher conductivity than Formulations A and B, whereas the difference between C and D was not statistically significant.

Tukey法で調整した群間比較では、配合Dの伝導度が配合AおよびBより高く、CとDの差は統計的に有意ではありませんでした。

二元配置分散分析

There were significant main effects of temperature (F(2, 54) = 12.8, p < 0.001) and composition (F(1, 54) = 7.3, p = 0.009), as well as a temperature-by-composition interaction (F(2, 54) = 4.6, p = 0.014).

温度(F(2, 54)=12.8、p<0.001)と組成(F(1, 54)=7.3、p=0.009)の主効果に加え、温度と組成の交互作用も有意でした(F(2, 54)=4.6、p=0.014)。

  • The significant interaction indicates that the effect of temperature depended on composition.
    有意な交互作用は、温度の効果が組成によって異なることを示しています。
  • We therefore examined simple effects within each composition.
    そのため、各組成内で単純効果を調べました。

相関分析の結果を英語で書く

Pearsonの相関

Porosity was negatively correlated with compressive strength (Pearson’s r = -0.71, 95% CI, -0.84 to -0.49; p < 0.001; n = 42).

空隙率は圧縮強度と負の相関を示しました(Pearsonのr=−0.71、95%信頼区間:−0.84~−0.49、p<0.001、n=42)。

Spearmanの順位相関

Because the relationship was monotonic but nonlinear, we used Spearman’s rank correlation. The correlation was moderate (ρ = 0.56, 95% bootstrap CI, 0.29 to 0.74; p = 0.001).

関係が単調ではあるものの非線形だったため、Spearmanの順位相関を使用しました。相関は中程度でした(ρ=0.56、ブートストラップ95%信頼区間:0.29~0.74、p=0.001)。

回帰分析の結果を英語で書く

線形回帰

After adjustment for batch and specimen thickness, a 10-degree increase in curing temperature was associated with a 1.8-MPa increase in adhesion strength (β = 1.8 MPa per 10°C, 95% CI, 0.9 to 2.7; p < 0.001).

バッチと試料厚さを調整した後、硬化温度が10℃上昇すると接着強度が1.8 MPa増加する関連が認められました(β=10℃当たり1.8 MPa、95%信頼区間:0.9~2.7、p<0.001)。

  • The model explained 62% of the observed variance (adjusted R² = 0.62).
    モデルは観測された分散の62%を説明しました(調整済みR²=0.62)。
  • The residual diagnostics did not indicate a substantial violation of the model assumptions.
    残差診断では、モデル仮定の重大な逸脱は示されませんでした。
  • The coefficient represents an association conditional on the included covariates.
    この係数は、モデルに含めた共変量で条件づけた関連を表します。

ロジスティック回帰

Higher moisture content was associated with increased odds of coating failure (adjusted odds ratio = 1.42 per 1% increase, 95% CI, 1.15 to 1.76; p = 0.001).

含水率の上昇は、コーティング故障のオッズ増加と関連していました(1%上昇当たりの調整オッズ比=1.42、95%信頼区間:1.15~1.76、p=0.001)。

比率・カテゴリーデータの結果を書く

カイ二乗検定

The pass rate differed between the two processes (82% vs. 64%; χ²(1) = 5.12, p = 0.024). The absolute difference was 18 percentage points (95% CI, 3 to 33 percentage points).

合格率は2つのプロセス間で異なりました(82%対64%、χ²(1)=5.12、p=0.024)。絶対差は18ポイントでした(95%信頼区間:3~33ポイント)。

Fisherの正確確率検定

Because the expected cell counts were small, we used Fisher’s exact test. The observed difference was not statistically significant (p = 0.14), and the confidence interval was wide.

期待度数が小さかったため、Fisherの正確確率検定を使用しました。観測された差は統計的に有意ではなく(p=0.14)、信頼区間は広いものでした。

ノンパラメトリック検定の結果を書く

Mann–Whitney U検定

The response values were higher in Group A than in Group B (median, 14.2 vs. 10.8 units; Mann–Whitney U = 168, p = 0.018). The estimated median difference was 3.1 units (95% bootstrap CI, 0.6 to 5.4).

応答値はグループBよりグループAで高い値でした(中央値:14.2対10.8単位、Mann–Whitney U=168、p=0.018)。推定中央値差は3.1単位でした(ブートストラップ95%信頼区間:0.6~5.4)。

Wilcoxon符号付順位検定

The measurements increased after treatment (median change, 2.4 units; Wilcoxon signed-rank test, V = 87, p = 0.009).

測定値は処理後に増加しました(変化量の中央値:2.4単位、Wilcoxon符号付順位検定:V=87、p=0.009)。

ノンパラメトリック検定を使ったからといって、常に「中央値の差」を検定しているとは限りません。検定が比較する仮説と、併記する効果量を解析設計に合わせます。

多重比較を英語で報告する

  • P-values were adjusted for multiple comparisons using the Holm method.
    Holm法を用いて多重比較のp値を調整しました。
  • The false discovery rate was controlled using the Benjamini–Hochberg procedure.
    Benjamini–Hochberg法を用いて偽発見率を制御しました。
  • After Bonferroni correction, two of the five comparisons remained statistically significant.
    Bonferroni補正後、5つの比較のうち2つが統計的に有意なままでした。
  • The reported confidence intervals are simultaneous 95% intervals.
    報告した信頼区間は同時95%信頼区間です。
  • The analysis was exploratory, and the p-values were not adjusted for multiplicity.
    この解析は探索的であり、p値の多重性調整は行っていません。

調整を行わなかった場合も、その事実と探索的解析であることを明示します。

境界付近のp値をどう書くか

marginally significant、almost significant、approached significanceは、p=0.05の境界を過度に強調することがあります。p=0.049とp=0.051を正反対の結果として扱うより、推定値と信頼区間を示します。

The estimated mean difference was 2.1 MPa (95% CI, -0.1 to 4.3 MPa; p = 0.061). The interval includes both a negligible effect and an improvement that may be practically relevant.

推定平均差は2.1 MPaでした(95%信頼区間:−0.1~4.3 MPa、p=0.061)。この区間には、無視できる効果と実用上意味を持つ可能性のある改善の両方が含まれています。

図表で統計結果を説明する英語

エラーバーを説明する

  • Points represent group means, and error bars show 95% confidence intervals.
    点は各群の平均値、エラーバーは95%信頼区間を示します。
  • The box shows the interquartile range, and the center line indicates the median.
    箱は四分位範囲、中央線は中央値を示します。
  • Individual observations are shown to display the underlying variability.
    データのばらつきを示すため、個々の観測値も表示しています。

群間差を説明する

  • The estimated differences and their confidence intervals are shown in the forest plot.
    推定差と信頼区間をフォレストプロットに示しています。
  • Positive values favor the modified formulation.
    正の値は改良配合が優れていることを表します。
  • The vertical dashed line represents the null value.
    縦の破線は帰無値を表します。
  • The shaded region indicates the predefined range of practically negligible effects.
    網掛け部分は、事前に定めた実用上無視できる効果の範囲を示します。

グラフの一般的な説明表現については、英語でグラフを説明する表現集も参照してください。

材料研究の統計結果を英語で書く例

以下は説明用の架空例です。

配合と硬化温度の影響

We evaluated the effects of formulation and curing temperature on coating adhesion using a two-way ANOVA. Both formulation (F(2, 54) = 9.6, p < 0.001; partial η² = 0.26) and temperature (F(1, 54) = 14.1, p < 0.001; partial η² = 0.21) had statistically significant main effects. The formulation-by-temperature interaction was also significant (F(2, 54) = 4.2, p = 0.020), indicating that the effect of temperature differed among formulations.

At 140°C, Formulation C exceeded the reference by 3.8 MPa (95% CI, 1.4 to 6.2 MPa; Holm-adjusted p = 0.004). At 120°C, the estimated difference was 0.9 MPa (95% CI, -1.2 to 3.0 MPa; adjusted p = 0.62). Thus, the advantage of Formulation C was concentrated at the higher curing temperature.

日本語訳:二元配置分散分析を用いて、配合と硬化温度がコーティング接着強度に与える影響を評価しました。配合(F(2, 54)=9.6、p<0.001、偏η²=0.26)と温度(F(1, 54)=14.1、p<0.001、偏η²=0.21)の両方に統計的に有意な主効果が認められました。配合と温度の交互作用も有意であり(F(2, 54)=4.2、p=0.020)、温度の効果が配合によって異なることが示されました。

140℃では、配合Cが基準配合を3.8 MPa上回りました(95%信頼区間:1.4~6.2 MPa、Holm調整済みp=0.004)。120℃での推定差は0.9 MPaでした(95%信頼区間:−1.2~3.0 MPa、調整済みp=0.62)。したがって、配合Cの利点は高い硬化温度で主に認められました。

有意差と実用性を分けて述べる

Although the modified process produced a statistically significant reduction in cycle time (mean difference, -1.4 min; 95% CI, -2.0 to -0.8 min; p < 0.001), the improvement did not reach the predefined operational target of 3 min.

改良プロセスではサイクル時間が統計的に有意に短縮されましたが(平均差:−1.4分、95%信頼区間:−2.0~−0.8分、p<0.001)、事前に定めた運用上の目標である3分には達しませんでした。

非有意結果を適切に説明する

The estimated difference in durability was 6 cycles in favor of the modified material, but the confidence interval was wide and included no difference (95% CI, -4 to 16 cycles; p = 0.23). The present experiment therefore does not distinguish between a negligible effect and a potentially useful improvement.

耐久性の推定差は改良材料で6サイクル高い値でしたが、信頼区間は広く、差がない場合も含んでいました(95%信頼区間:−4~16サイクル、p=0.23)。したがって今回の実験では、無視できる効果と実用的な改善の可能性を区別できません。

Resultsセクションで使える文章構成

1.最初に記述統計を示す

Mean conductivity was 1.42 ± 0.18 S cm⁻¹ in the modified group and 1.21 ± 0.16 S cm⁻¹ in the reference group.

2.推定差と信頼区間を示す

The estimated mean difference was 0.21 S cm⁻¹ (95% CI, 0.08 to 0.34 S cm⁻¹).

3.検定結果と効果量を加える

The difference was statistically significant (Welch’s t(17.6) = 3.39, p = 0.003), with a standardized mean difference of d = 1.07.

4.限定された意味を述べる

The confidence interval remained above the predefined minimum relevant difference of 0.05 S cm⁻¹, supporting practical relevance within the tested composition range.

Abstractで簡潔に報告する

The modified formulation increased mean adhesion strength by 3.7 MPa compared with the reference (95% CI, 1.8 to 5.6 MPa; p = 0.001; Cohen’s d = 1.02) and exceeded the predefined practical threshold.

要旨では、主要評価項目について群別の代表値、効果の推定値、信頼区間を優先し、必要に応じてp値を加えます。

質疑応答で統計結果を説明する

「p値が小さいので効果も大きいですか」

Not necessarily. The p-value reflects the compatibility of the data with the null model, not the magnitude of the effect. The effect size was 1.2 MPa, and the 95% confidence interval ranged from 0.8 to 1.6 MPa.

必ずしもそうではありません。p値は帰無モデルとデータの整合性を反映するもので、効果の大きさそのものではありません。効果量は1.2 MPaで、95%信頼区間は0.8~1.6 MPaでした。

「有意差がないので、同じですか」

The test did not detect a statistically significant difference, but that does not establish equivalence. The confidence interval still includes differences that may be practically relevant.

検定では統計的に有意な差を検出しませんでしたが、それは同等性を確立するものではありません。信頼区間には、実用上意味を持つ可能性のある差がまだ含まれています。

「なぜ効果量を示すのですか」

The effect size shows the magnitude of the difference, whereas the p-value alone does not. Reporting both the effect estimate and its confidence interval allows us to assess practical importance and precision.

効果量は差の大きさを示しますが、p値だけでは大きさは分かりません。効果の推定値と信頼区間を報告することで、実用上の重要性と推定精度を評価できます。

「サンプル数は十分ですか」

The sample size was determined for the primary outcome using a minimum relevant difference of 2 MPa. However, the study was not powered to estimate subgroup effects precisely, so those analyses should be considered exploratory.

主要評価項目について、実質的に意味のある最小差を2 MPaとしてサンプル数を決定しました。ただし、サブグループ効果を精密に推定する検出力はないため、それらの解析は探索的と位置づけています。

よくある英語表現の修正

避けたい表現 より正確な表現
There was a significant difference (p < 0.05). The mean difference was 3.7 MPa (95% CI, 1.8 to 5.6 MPa; p = 0.001).
The null hypothesis was accepted. We did not reject the null hypothesis.
There was no difference. The difference was not statistically significant, and the estimate was imprecise.
The p-value proves the effect. The data provided evidence against the null model under the stated assumptions.
p = 0.000 p < 0.001
The result was nearly significant. The estimate was 2.1 MPa (95% CI, -0.1 to 4.3 MPa; p = 0.061).
The correlation caused the improvement. The variables were associated; causal interpretation requires additional evidence.

数値・記号の書き方

  • p値は通常、pをイタリック体にします。投稿先の規程に従ってください。
  • p値は必要な精度で示し、0.000とは書きません。
  • 信頼区間は、95% CI, 1.2 to 3.4または95% CI [1.2, 3.4]など、投稿先の形式に合わせます。
  • 平均値には、SD、SE、CIのどれを付けたか明記します。
  • percentage pointsとpercentを区別します。50%から60%への増加は10 percentage points、相対的には20 percentです。
  • 単位は数値とともに示し、比較文でも意味が曖昧にならないようにします。

統計結果を書くときのチェックリスト

  • 各群のサンプル数と記述統計を示したか
  • 主要な効果を元の単位で示したか
  • 信頼区間を示したか
  • p値を可能な範囲で正確に示したか
  • 検定統計量と自由度を必要に応じて示したか
  • 効果量の種類を明記したか
  • 統計的有意性と実務上の重要性を分けたか
  • 非有意を「差がない証明」としていないか
  • 多重比較の調整方法を示したか
  • 探索的解析と確認的解析を区別したか
  • 解析の前提と主要な不確かさを説明したか

そのまま使える統計結果テンプレート

2群比較

The mean [outcome] was [value] in Group A and [value] in Group B. The mean difference was [estimate] (95% CI, [lower] to [upper]; [test statistic], p = [value]; [effect size]). This difference [did/did not] exceed the predefined practical threshold of [value].

相関

[Variable A] was [positively/negatively] correlated with [Variable B] ([Pearson’s r/Spearman’s ρ] = [value], 95% CI, [lower] to [upper]; p = [value]; n = [number]). This association should be interpreted within [scope or limitation].

回帰

After adjustment for [covariates], a [unit] increase in [predictor] was associated with a [estimate]-unit [increase/decrease] in [outcome] (β = [value], 95% CI, [lower] to [upper]; p = [value]). The model explained [value]% of the observed variance.

非有意結果

The estimated difference was [value] (95% CI, [lower] to [upper]; p = [value]). Although the result was not statistically significant, the interval [included/excluded] effects that would be practically meaningful.

関連する英語・統計記事

まとめ

統計解析結果を英語で書くときは、p値だけで有意・非有意を述べるのではなく、推定値、効果量、信頼区間、p値、実務的な意味を一続きで示します。

基本形は次の一文です。

The estimated difference was [effect size] (95% CI, [lower] to [upper]; p = [value]), and the effect [did/did not] exceed the predefined threshold for practical relevance.

推定差は[効果量]でした(95%信頼区間:[下限]~[上限]、p=[値])。この効果は、事前に定めた実用上の閾値を[上回りました/上回りませんでした]。

この形式なら、差の方向と大きさ、推定の精度、仮説検定の結果、研究上の意味を短い文章で伝えられます。非有意の場合も、信頼区間を用いることで「効果がない」と断定せず、今回のデータがどの範囲の効果と整合するかを説明できます。

参考文献

  1. Wasserstein, R. L. and Lazar, N. A., “The ASA’s Statement on p-Values: Context, Process, and Purpose,” The American Statistician, 70(2), 129–133 (2016).
  2. Lang, T. A. and Altman, D. G., “Basic Statistical Reporting for Articles Published in Biomedical Journals: The SAMPL Guidelines.”
  3. Ranstam, J. et al., “On reporting and interpreting statistical significance and p values in medical research,” BMC Medicine, 19, 38 (2021).
  4. Wasserstein, R. L., Schirm, A. L. and Lazar, N. A., “Moving to a World Beyond p < 0.05,” The American Statistician, 73(sup1), 1–19 (2019).
  5. PLOS Complex Systems, “Best Practices in Research Reporting.”

コメントする

メールアドレスが公開されることはありません。 ※ が付いている欄は必須項目です

CAPTCHA


上部へスクロール