The training is taking forever after 2 Hr it haven't even reached halfway yet but from the way the model is learning i have assessed the need to lower the learning rate and i might also need to augment the dataset more so it can simulate more noise the actual real world scenario might bring for the next training. am also planning to test the new transformer concept in here 😂 just to get my hand dirty and know more about it even though using it here is an overkill my planed augmentation will boost the total dataset well above 300k a simple CNN-Attention block or a lightweight Transformer encoder is feasible