Back to models
Transformer

SFA

Symbolic Fourier Approximation (SFA) Transformer.

Overview: for each series:

run a sliding window across the series for each window

shorten the series with DFT discretise the shortened series into bins set by MFC form a word from these discrete values

by default SFA produces a single word per series (window_size=0) if a window is used, it forms a histogram of counts of words.

Quickstart

python
from sktime.transformations.dictionary_based import SFA

estimator = SFA(word_length=8, alphabet_size=4, window_size=12, norm=False, binning_method='equi-depth', anova=False, bigrams=False, skip_grams=False, remove_repeat_words=False, levels=1, lower_bounding=True, save_words=False, keep_binning_dft=False, return_pandas_data_series=False, use_fallback_dft=False, typed_dict=False, n_jobs=1)

Tags

Capabilities

  • Unequal-length series
  • Multivariate: Not supported
  • Inverse transform: Not supported
  • Missing values: Not supported
  • Removes missing values: Not supported
  • Equalizes series length: Not supported

Properties

Input typescitype:transform-input
Series
Output typescitype:transform-output
Series
Label typescitype:transform-labels
None
Fit is emptyfit_is_empty
Yes
Keeps the time indextransform-returns-same-time-index
No
Requires Xrequires_X
Yes
Requires yrequires_y
Yes
X and y need the same indexX-y-must-have-same-index
No

Parameters(13)

word_length: int, default = 8
length of word to shorten window to (using PAA)
alphabet_size: int, default = 4
number of values to discretise each value to
window_size: int, default = 12
size of window for sliding. Input series length for whole series transform
norm: boolean, default = False
mean normalise words by dropping first fourier coefficient
binning_method: {“equi-depth”, “equi-width”, “information-gain”, “kmeans”},

default=”equi-depth”

the binning method used to derive the breakpoints.

anova: boolean, default = False
If True, the Fourier coefficient selection is done via a one-way ANOVA test. If False, the first Fourier coefficients are selected. Only applicable if labels are given
bigrams: boolean, default = False
whether to create bigrams of SFA words
skip_grams: boolean, default = False
whether to create skip-grams of SFA words
remove_repeat_words: boolean, default = False
whether to use numerosity reduction (default False)
levels: int, default = 1
Number of spatial pyramid levels
save_words: boolean, default = False
whether to save the words generated for each series (default False)
return_pandas_data_series: boolean, default = False
set to true to return Pandas Series as a result of transform. setting to true reduces speed significantly but is required for automatic test.
n_jobs: int, optional, default = 1

The number of jobs to run in parallel for both transform. -1 means using all processors.

References

[1]

Schäfer, Patrick, and Mikael Högqvist. “SFA: a symbolic fourier approximation

and index for similarity search in high dimensional datasets.” Proceedings of the 15th international conference on extending database technology. 2012.