来自数据框的神经网络LSTM输入形状

以下是设置时间序列数据以训练LSTM的示例。该模型输出是胡说八道，因为我仅将其设置为演示如何构建模型。

import pandas as pd
import numpy as np
# Get some time series data
df = pd.read_csv("https://raw.githubusercontent.com/plotly/datasets/master/timeseries.csv")
df.head()

时间序列数据帧：

Date      A       B       C      D      E      F      G
0   2008-03-18  24.68  164.93  114.73  26.27  19.21  28.87  63.44
1   2008-03-19  24.18  164.89  114.75  26.22  19.07  27.76  59.98
2   2008-03-20  23.99  164.63  115.04  25.78  19.01  27.04  59.61
3   2008-03-25  24.14  163.92  114.85  27.41  19.61  27.84  59.41
4   2008-03-26  24.44  163.45  114.84  26.86  19.53  28.02  60.09

您可以将输入输入构建为向量，然后使用pandas.cumsum()函数构建时间序列的序列：

# Put your inputs into a single list
df['single_input_vector'] = df[input_cols].apply(tuple, axis=1).apply(list)
# Double-encapsulate list so that you can sum it in the next step and keep time steps as separate elements
df['single_input_vector'] = df.single_input_vector.apply(lambda x: [list(x)])
# Use .cumsum() to include prevIoUs row vectors in the current row list of vectors
df['cumulative_input_vectors'] = df.single_input_vector.cumsum()

可以类似的方式设置输出，但是它将是单个向量而不是序列：

# If your output is multi-dimensional, you need to capture those dimensions in one object
# If your output is a single dimension, this step may be unnecessary
df['output_vector'] = df[output_cols].apply(tuple, axis=1).apply(list)

输入序列必须具有相同的长度才能在模型中运行，因此您需要将其填充为累积向量的最大长度：

# Pad your sequences so they are the same length
from keras.preprocessing.sequence import pad_sequences

max_sequence_length = df.cumulative_input_vectors.apply(len).max()
# Save it as a list   
padded_sequences = pad_sequences(df.cumulative_input_vectors.tolist(), max_sequence_length).tolist()
df['padded_input_vectors'] = pd.Series(padded_sequences).apply(np.asarray)

训练数据可以从数据框中提取并放入numpy数组中。

您可以使用hstack和reshape来构建3D输入数组。

# Extract your training data
X_train_init = np.asarray(df.padded_input_vectors)
# Use hstack to and reshape to make the inputs a 3d vector
X_train = np.hstack(X_train_init).reshape(len(df),max_sequence_length,len(input_cols))
y_train = np.hstack(np.asarray(df.output_vector)).reshape(len(df),len(output_cols))

为了证明这一点：

>>> print(X_train_init.shape)
(11,)
>>> print(X_train.shape)
(11, 11, 6)
>>> print(X_train == X_train_init)
False

获得训练数据后，您可以定义输入层和输出层的尺寸。

# Get your input dimensions
# Input length is the length for one input sequence (i.e. the number of rows for your sample)
# Input dim is the number of dimensions in one input vector (i.e. number of input columns)
input_length = X_train.shape[1]
input_dim = X_train.shape[2]
# Output dimensions is the shape of a single output vector
# In this case it's just 1, but it Could be more
output_dim = len(y_train[0])

建立模型：

from keras.models import Model, Sequential
from keras.layers import LSTM, Dense

# Build the model
model = Sequential()

# I arbitrarily picked the output dimensions as 4
model.add(LSTM(4, input_dim = input_dim, input_length = input_length))
# The max output value is > 1 so relu is used as final activation.
model.add(Dense(output_dim, activation='relu'))

model.compile(loss='mean_squared_error',
              optimizer='sgd',
              metrics=['accuracy'])

最后，您可以训练模型并将训练日志保存为历史记录：

# Set batch_size to 7 to show that it doesn't have to be a factor or multiple of your sample size
history = model.fit(X_train, y_train,
              batch_size=7, nb_epoch=3,
              verbose = 1)

输出：

Epoch 1/3
11/11 [==============================] - 0s - loss: 3498.5756 - acc: 0.0000e+00     
Epoch 2/3
11/11 [==============================] - 0s - loss: 3498.5755 - acc: 0.0000e+00     
Epoch 3/3
11/11 [==============================] - 0s - loss: 3498.5757 - acc: 0.0000e+00

而已。使用model.predict(X)whereX与为X_train从模型进行预测时使用相同的格式（样本数量除外）。

其他 2022/1/1 18:33:12 有492人围观

撰写回答

你尚未登录，登录后可以

和开发者交流问题的细节

关注并接收问题和回答的更新提醒

参与内容的编辑和改进，让解决方法与时俱进

请先登录

来自数据框的神经网络LSTM输入形状

撰写回答

推荐问题

模块构建失败（来自./node_modules/babel-loader/lib/index.js）：错误：找不到模块“ babel-preset- react”

来自SelectList的ASP.NET MVC下拉列表

来自jsp：include的response.sendRedirect（）是否被忽略？

来自其他容器的Docker mongo映像``连接被拒绝''

来自编号的MySQL MONTHNAME（）

来自Google Takeout的jpg批量加入json和jpg

linux / gcc：来自C / C ++程序的ldd功能

包含来自容器的日志的日志文件在哪里？

在最近的邻居算法中“来自不同的顶点链”是什么意思？

如何使用selenium和scrapy来自动化该过程？

来自变量的mysql字段名称

来自MySQL中多个表的COUNT（*）

来自pip import main的<module>中文件“ / usr / bin / pip”，第9行，ImportError：无法导入名称main

来自JSP的可编辑Word文档

如何正确处理来自ListenableFuture番石榴的异常？

如何序列化/反序列化为字典`来自自定义XML而不使用XElement？

PHP-如何最好地确定当前调用是来自CLI还是Web服务器？

Spring Boot / Spring Security正在忽略来自JavaScript的/ login调用

实体框架代码优先-来自同一表的两个外键

编辑来自Google服务帐户的Google日历事件：403

分类汇总

您的鼓励是对我最大的支持