[`dplyr::filter()`](https://dplyr.tidyverse.org/reference/filter.html) は、stringr パッケージの文字列関数を組み合わせることで、部分一致・前方一致・後方一致による柔軟な文字列マッチングが可能である。 ## 部分一致 (`str_detect()`) [`stringr::str_detect()`](https://stringr.tidyverse.org/reference/str_detect.html) は、文字列の任意の位置に指定したパターンが含まれているかを検索する場合に使用する。 ```r babynames |> filter(str_detect(name, r"([L|l]ee)")) ``` ``` # A tibble: 21,717 × 5 year sex name n prop <dbl> <chr> <chr> <int> <dbl> 1 1880 F Lee 28 0.000287 2 1880 F Kathleen 21 0.000215 3 1880 F Paralee 7 0.0000717 4 1880 F Aileen 6 0.0000615 5 1880 F Rosalee 5 0.0000512 6 1880 M Lee 361 0.00305 7 1881 F Lee 39 0.000395 8 1881 F Kathleen 13 0.000132 9 1881 F Paralee 13 0.000132 10 1881 M Lee 342 0.00316 # ℹ 21,707 more rows # ℹ Use `print(n = ...)` to see more rows ``` ## 前方一致 (`str_starts()`) [`stringr::str_starts()`](https://stringr.tidyverse.org/reference/str_starts.html) は、文字列の先頭に指定したパターンがあるかを検索する場合に使用する。 ```r babynames |> filter(str_starts(name, "Alex")) ``` ``` # A tibble: 3,659 × 5 year sex name n prop <dbl> <chr> <chr> <int> <dbl> 1 1880 M Alexander 211 0.00178 2 1880 M Alex 147 0.00124 3 1881 M Alexander 209 0.00193 4 1881 M Alex 114 0.00105 5 1882 F Alexina 5 0.0000432 6 1882 M Alexander 225 0.00184 7 1882 M Alex 172 0.00141 8 1882 M Alexis 6 0.0000492 9 1883 M Alexander 187 0.00166 10 1883 M Alex 120 0.00107 # ℹ 3,649 more rows # ℹ Use `print(n = ...)` to see more rows ``` ## 後方一致 (`str_ends()`) [`stringr::str_ends()`](https://stringr.tidyverse.org/reference/str_starts.html) は、文字列の末尾に指定したパターンがあるかを検索する場合に使用する。 ```r babynames |> filter(str_ends(name, "th")) ``` ``` # A tibble: 10,889 × 5 year sex name n prop <dbl> <chr> <chr> <int> <dbl> 1 1880 F Elizabeth 1939 0.0199 2 1880 F Edith 768 0.00787 3 1880 F Ruth 234 0.00240 4 1880 F Elisabeth 24 0.000246 5 1880 F Elizebeth 18 0.000184 6 1880 F Edyth 6 0.0000615 7 1880 F Faith 6 0.0000615 8 1880 F Judith 5 0.0000512 9 1880 M Kenneth 32 0.000270 10 1880 M Seth 25 0.000211 # ℹ 10,879 more rows # ℹ Use `print(n = ...)` to see more rows ``` ## 関連ノート - **[前提]** [[dplyr の filter() で論理演算子を使用する]]: 文字列マッチング関数はこの `filter()` の中で組み合わせて使うため、先にこちらを押さえると組み合わせが理解しやすい